arXiv:2510.08460cs.CL2025-10被引 17

构建多任务分歧感知基准,推动AI理解人类判断差异。

LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task

  • 扩展至4个数据集,支持分类与序数标注的分歧建模。
  • 引入软标签与视角主义双评估范式,超越传统交叉熵指标。
  • 适合关注人机判断差异、鲁棒性评估的研究者参考。

许多研究者认为,AI模型应具备识别人类判断变异与分歧的能力,并据此进行训练与评估。为推进这一理念,学习分歧共享任务(LEWIDI)系列旨在提升相关数据集的可获取性并发展评估方法。第三届任务在此基础上,将LEWIDI基准扩展至四个数据集:复述识别、讽刺检测、反语检测和自然语言推理,标注体系不仅包含以往的分类判断,还引入了序数判断。另一创新在于采用两种互补的评估范式:软标签法(预测群体判断分布)与视角主义法(预测个体标注者观点)。关键的是,我们摒弃标准交叉熵等指标,测试了适用于两类范式的新型评估方法。任务吸引了多样化的参与,结果揭示了建模判断变异方法的优势与局限。这些贡献强化了LEWIDI框架,提供了新资源、基准与发现,助力分歧感知技术的发展。

原文摘要 · Abstract (English)

Many researchers have reached the conclusion that AI models should be trained to be aware of the possibility of variation and disagreement in human judgments, and evaluated as per their ability to recognize such variation. The LEWIDI series of shared tasks on Learning With Disagreements was established to promote this approach to training and evaluating AI models, by making suitable datasets more accessible and by developing evaluation methods. The third edition of the task builds on this goal by extending the LEWIDI benchmark to four datasets spanning paraphrase identification, irony detection, sarcasm detection, and natural language inference, with labeling schemes that include not only categorical judgments as in previous editions, but ordinal judgments as well. Another novelty is that we adopt two complementary paradigms to evaluate disagreement-aware systems: the soft-label approach, in which models predict population-level distributions of judgments, and the perspectivist approach, in which models predict the interpretations of individual annotators. Crucially, we moved beyond standard metrics such as cross-entropy, and tested new evaluation metrics for the two paradigms. The task attracted diverse participation, and the results provide insights into the strengths and limitations of methods to modeling variation. Together, these contributions strengthen LEWIDI as a framework and provide new resources, benchmarks, and findings to support the development of disagreement-aware technologies.

分歧感知多任务评估方法自然语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。