Beyond edge addition
| Item Type: | Dataset |
|---|---|
| Title: | Beyond edge addition |
| Alternative Title: | Beyond edge addition: A dataset for information extraction incorporating new instances, types, and relations |
| Date: | 2025 |
| Creator: |
Hertling, Sven ORCID: 0000-0003-0333-5888 ; Möller, Cedric ; Mihindukulasooriya, Nandana ; Usbeck, Ricardo
|
| Divisions: | School of Business Informatics and Mathematics > Sonstige - Fakultät für Wirtschaftsinformatik und Wirtschaftsmathematik School of Business Informatics and Mathematics > Data Science (Paulheim 2018-) |
| DDC Classification: |
004 Computer science, internet |
|---|---|
| Abstract: | Information extraction (IE) is the task of converting natural language text into structured triples comprising a subject, predicate, and object. Existing IE datasets often operate under the assumption that all entities (instances, properties, and classes) are already defined within a knowledge graph (KG), focusing solely on discovering the relationships between them. However, this assumption does not align with real-world scenarios, where many entities and relationships may be missing in the KG. Additionally, most current datasets do not provide a snapshot of the accompanying knowledge graph, leading to inconsistencies in evaluation, as different systems may rely on different KG versions with varying degrees of completeness and labelling support. Such inconsistencies undermine fair benchmarking and reproducibility. In this paper, we introduce a novel information extraction dataset specifically designed to better reflect realistic KG incompleteness. Our dataset includes 20% missing classes and instances, along with 5% missing relations, requiring systems to not only add new links (edges) but also propose new instances, classes, and relations. To ensure reproducibility and prevent leakage from pre-trained language models, we provide a heavily modified version of Wikidata where background knowledge cannot be exploited to trivially infer triples. This resource supports a more robust and comparable evaluation of IE systems in settings closer to real-world applications. We further present a strong baseline that employs large language models for extraction and disambiguation tasks, as well as encoder-based retrieval, to integrate the background knowledge graph. It operates in several iterations, initially identifying a first set of triples, then progressively refining them by referencing the KG and generating new entities if necessary. (English) |
| External Identifier for Data: | https://doi.org/10.5281/zenodo.17708383 |
| URL: | https://madata.bib.uni-mannheim.de/1050/ |
|---|---|
| Access (Controlled): | Only Metadata |
| License (Controlled): | No license information available, not specified, or non-standard license |
Full text not available from this repository.
| Date Deposited: | 24 Jun 2026 15:32 |
|---|---|
| Last Modified: | 24 Jun 2026 15:32 |
You have found an error? Please let us know about your desired correction here: E-Mail
Actions (login required)
![]() |
View Item |

