arXiv:2502.13853cs.CL2025-02NAACL被引 14

首个支持多答案与人工分歧的谬误检测数据集,提升评估真实性。

Fine-grained Fallacy Detection with Human Label Variation

  • 构建含11000+标注的意大利语社交媒体谬误数据集,支持20类谬误
  • 提出新评估框架,兼容多可靠答案与部分重叠标注
  • 适合研究谬误识别与人类标注差异的学者使用

我们提出Faina,首个接纳多种合理答案与自然分歧的谬误检测数据集。该数据集包含超过11,000个跨度级标注,覆盖意大利语社交媒体帖子中关于移民、气候变化和公共健康的20类谬误,由两名专家在多轮讨论下完成标注。通过深入的标注研究,我们减少了标注误差,同时保留了人类标注的自然变异。此外,我们设计了一种超越‘单一真实答案’评估的新框架,可同时处理多个同等可靠的测试集,并考虑任务特性,如部分跨度匹配、类别重叠及标注错误严重程度差异。在四种谬误检测设置下的实验表明,多任务与多标签的Transformer方法在所有场景中均为强基线。我们公开数据、代码与标注指南,以推动谬误检测与人类标注差异研究的发展。

原文摘要 · Abstract (English)

We introduce Faina, the first dataset for fallacy detection that embraces multiple plausible answers and natural disagreement. Faina includes over 11K span-level annotations with overlaps across 20 fallacy types on social media posts in Italian about migration, climate change, and public health given by two expert annotators. Through an extensive annotation study that allowed discussion over multiple rounds, we minimize annotation errors whilst keeping signals of human label variation. Moreover, we devise a framework that goes beyond "single ground truth" evaluation and simultaneously accounts for multiple (equally reliable) test sets and the peculiarities of the task, i.e., partial span matches, overlaps, and the varying severity of labeling errors. Our experiments across four fallacy detection setups show that multi-task and multi-label transformer-based approaches are strong baselines across all settings. We release our data, code, and annotation guidelines to foster research on fallacy detection and human label variation more broadly.

谬误检测多答案标注人类差异数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。