Is it novel and why? Fine-grained patent novelty prediction based on passage retrieval
- Novelty assessment is a critical yet complex task in the examination process for patent acceptance, requiring examiners to determine whether an invention is disclosed in a prior art document. The process involves intricate matching between specific features of a patent claim and passages in the prior art. While prior work has approached novelty prediction primarily as a binary classification task at the claim level, we argue that this formulation is susceptible to spurious correlations and lacks the granularity required for practical application. In this work, we introduce FiNE-Patents (Fine-grained Novelty Examination of Patents), a novel dataset comprising 3,658 first patent claims annotated with fine-grained, feature-level prior art references extracted from European Search Opinion (ESOP) documents. We propose shifting the evaluation paradigm from simple binary classification to a joint retrieval and abstract reasoning task at the feature level, requiring models to identify specificNovelty assessment is a critical yet complex task in the examination process for patent acceptance, requiring examiners to determine whether an invention is disclosed in a prior art document. The process involves intricate matching between specific features of a patent claim and passages in the prior art. While prior work has approached novelty prediction primarily as a binary classification task at the claim level, we argue that this formulation is susceptible to spurious correlations and lacks the granularity required for practical application. In this work, we introduce FiNE-Patents (Fine-grained Novelty Examination of Patents), a novel dataset comprising 3,658 first patent claims annotated with fine-grained, feature-level prior art references extracted from European Search Opinion (ESOP) documents. We propose shifting the evaluation paradigm from simple binary classification to a joint retrieval and abstract reasoning task at the feature level, requiring models to identify specific passages from a prior art document that disclose individual claim features, and to identify which features of a claim make it novel. We implement and evaluate LLM-based workflows that decompose claims into features, analyze each feature against prior art, and finally derive a claim-level novelty prediction. Our experiments demonstrate that these workflows outperform embedding-based baselines on passage retrieval and novel feature identification. Furthermore, we show that unlike trained classifiers, LLMs are robust against spurious correlations present in the claim-level novelty classification task. We release the dataset and code to foster further research into transparent and granular patent analysis.…


| Author: | Valentin KnappichORCiD, Anna Hätty, Simon Razniewski, Annemarie FriedrichORCiDGND |
|---|---|
| URN: | urn:nbn:de:bvb:384-opus4-1322891 |
| Frontdoor URL | https://opus.bibliothek.uni-augsburg.de/opus4/132289 |
| ISBN: | 979-8-4007-2599-9OPAC |
| Parent Title (English): | SIGIR '26: proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, Melbourne, VIC, Australia, July 20-24, 2026 |
| Publisher: | ACM |
| Place of publication: | New York, NY |
| Editor: | Alistair Moffat, Falk Scholer, Hannah Bast, Marc Najork, Min Zhang |
| Type: | Conference Proceeding |
| Language: | English |
| Year of first Publication: | 2026 |
| Publishing Institution: | Universität Augsburg |
| Release Date: | 2026/07/29 |
| First Page: | 845 |
| Last Page: | 856 |
| DOI: | https://doi.org/10.1145/3805712.3809576 |
| Institutes: | Fakultät für Angewandte Informatik |
| Fakultät für Angewandte Informatik / Institut für Informatik | |
| Fakultät für Angewandte Informatik / Institut für Informatik / Lehrstuhl für Computerlinguistik | |
| Dewey Decimal Classification: | 0 Informatik, Informationswissenschaft, allgemeine Werke / 00 Informatik, Wissen, Systeme / 004 Datenverarbeitung; Informatik |
| Licence (German): | CC-BY 4.0: Creative Commons: Namensnennung |



