提出新评估方法,解决捷克斯洛伐克新闻评论中长文本论断抽取的评测难题。
Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak
- 通过匹配和相似度评分实现多组论断间的最优对齐
- 在新收集的双语评论数据集上验证,发现现有方法存在明显不足
- 适用于模型性能评估与标注者一致性分析,适合事实核查研究者
文档级论断抽取仍是事实核查领域的开放挑战,相关评估方法研究有限。本文探索如何对同一源文档的两组论断进行对齐,并通过对齐得分计算其相似性,研究最佳对齐策略与评估方式,旨在构建可靠的评估框架。该方法可比较模型抽取与人工标注的论断集,既可用于评估模型提取性能,也可作为标注者间一致性的衡量指标。实验基于新收集的数据集,涵盖捷克与斯洛伐克新闻文章下的评论,这些领域因非正式语言、强烈地方语境及两种密切关联语言的细微差异而更具挑战性。结果揭示了现有评估方法在文档级论断抽取中的局限性,强调需要更先进的方法来准确捕捉语义相似性,并有效评估论断的关键属性,如原子性、可验证性和去上下文化程度。
原文摘要 · Abstract (English)
Document-level claim extraction remains an open challenge in the field of fact-checking, and subsequently, methods for evaluating extracted claims have received limited attention. In this work, we explore approaches to aligning two sets of claims pertaining to the same source document and computing their similarity through an alignment score. We investigate techniques to identify the best possible alignment and evaluation method between claim sets, with the aim of providing a reliable evaluation framework. Our approach enables comparison between model-extracted and human-annotated claim sets, serving as a metric for assessing the extraction performance of models and also as a possible measure of inter-annotator agreement. We conduct experiments on newly collected dataset-claims extracted from comments under Czech and Slovak news articles-domains that pose additional challenges due to the informal language, strong local context, and subtleties of these closely related languages. The results draw attention to the limitations of current evaluation approaches when applied to document-level claim extraction and highlight the need for more advanced methods-ones able to correctly capture semantic similarity and evaluate essential claim properties such as atomicity, checkworthiness, and decontextualization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。