KIT | KIT-Bibliothek | Impressum | Datenschutz

Dataset of "Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents"

Hey, Tobias ORCID iD icon 1; Keim, Jan ORCID iD icon 2; Tichy, Walter F. ORCID iD icon 1
1 Institut für Programmstrukturen und Datenorganisation (IPD), Karlsruher Institut für Technologie (KIT)
2 Institut für Informationssicherheit und Verlässlichkeit (KASTEL), Karlsruher Institut für Technologie (KIT)

Abstract:

This is the dataset used in the paper "Knowledge-based Sense Disambiguation of Multiword Expressions in Requirements Documents" at AIRE'21 In this paper, we explore the use of a multiword expression detection in combination with a knowledge-based word sense disambiguation to disambiguate expressions in requirements documents. The dataset comprises a gold standard for multiword expression detection and sense disambiguation for Wikipedia and WordNet 3.1. It covers 18 projects: CM1, EBT and GANTT as well as the 15 projects of the NFR dataset. <strong>File format</strong> We use a tab-separated version of the DiMSUM file format and extended it with sense information. The nine original DiMSUM tab-separated columns: 1. token offset 2. word 3. lowercase lemma 4. POS 5. MWE tag 6. offset of parent token (i.e. previous token in the same MWE), if applicable 7. strength level encoded in the tag, if applicable. Currently not used 8. supersense label, Currently not used 9. sentence ID and the two further columns for sense information: 10. Wikipedia article name 11. WordNet 3.1 synset The last two columns might end with .1 or .0 indicating that the sense is a fully applicable or partial sense of a multiword expression. ... mehr


Download
Originalveröffentlichung
DOI: 10.5281/zenodo.5167247
Zugehörige Institution(en) am KIT Institut für Informationssicherheit und Verlässlichkeit (KASTEL)
Institut für Programmstrukturen und Datenorganisation (IPD)
Publikationstyp Forschungsdaten
Publikationsdatum 06.08.2021
Identifikator KITopen-ID: 1000179726
Lizenz Creative Commons Namensnennung 4.0 International
Schlagwörter Multiword Expressions, Word Sense Disambiguation, Requirements Engineering
Art der Forschungsdaten Dataset
Nachgewiesen in OpenAlex
KIT – Die Universität in der Helmholtz-Gemeinschaft
KITopen Landing Page