构建两个捷克语语料库,揭示人类标注者在指代与话语关系上的差异。
Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations
- 创建双语料库,多人并行标注指代与话语关系。
- 标注一致率约60-65%,模型分歧处人类也更难判断。
- 附注释揭示个体理解差异与阅读策略差异。
先前研究显示,不同个体对文本连贯性的理解存在显著差异。为深入探究此现象,我们构建了两个包含捷克语文本多标注的语料库,并附有标注者对其选择的解释。第一个语料库包含1,024个语境,由三位标注者并行标注,涵盖代词、完整名词短语及回指副词等各类语法-语义类别中的指代识别差异。第二个语料库包含512个语境,由五位标注者并行标注,聚焦于属性性与非属性性结构中的话语关系识别。两个语料库的标注者间一致性均达到约60-65%。在指代标注中,当自动指代消解模型出现分歧时,人类标注者的一致性也更低,表明这些例子对人类而言更具挑战或歧义。标注者的注释进一步揭示了理解差异、信心水平不一及个体阅读策略的多样性。
原文摘要 · Abstract (English)
As previous research on annotator disagreement in discourse phenomena has shown, understanding text coherence varies considerably from one individual to another. To explore this phenomenon, we created two corpora with multiple annotations of Czech texts, accompanied by annotators' explanations of their choices. The first corpus consists of 1,024 contexts annotated in parallel by three annotators. It captures differences in the identification of coreference across various text types and grammatical-semantic categories, including pronouns, full noun phrases, and anaphoric adverbials. The second corpus comprises 512 contexts, annotated in parallel by five annotators, and focuses on identifying discourse relations in attributive and non-attributive constructions. Both corpora achieve a comparable inter-annotator agreement of approximately 60-65%. For coreference annotation, agreement tends to be lower in cases where automatic coreference resolution models disagree, suggesting that when the models disagree, the examples tend to be more difficult or ambiguous for human annotators to interpret. The annotators' comments, both for coreference and discourse relations, further reveal differences in interpretation, varying levels of confidence in text understanding, and individual reading strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。