arXiv:2507.15114cs.CL2025-07中稿 · Workshop on Perspe…被引 4

让自然语言推理模型先识别语义模糊,再做判断,更贴近人类理解。

From Disagreement to Understanding: The Case for Ambiguity Detection in NLI

  • 在推理前加入模糊性检测与分类,提升模型鲁棒性。
  • 提出统一的模糊类型分类框架,支持细粒度分析。
  • 适合追求可解释性和人类对齐的NLI研究者使用。

本文主张,自然语言推理(NLI)中的标注分歧并非噪声,而是由前提或假设中语义模糊引发的有意义差异。尽管标注指南不明确和标注者行为会带来变异性,但内容层面的模糊性提供了独立于流程的、反映人类多元视角的信号。我们呼吁构建以模糊性感知为核心的NLI系统:首先识别模糊输入对,分类其类型,再进行推理。为此,我们提出一个在推理前集成模糊性检测与分类的框架,并构建了一个整合现有分类体系的统一术语体系,通过实例展示关键子类型,推动针对性检测方法的发展。当前虽缺乏明确标注模糊性及子类型的资源,但这一空白带来了新机遇:通过开发新的标注数据集并探索无监督模糊检测方法,可实现更鲁棒、可解释且与人类理解一致的NLI系统。

原文摘要 · Abstract (English)

This position paper argues that annotation disagreement in Natural Language Inference (NLI) is not mere noise but often reflects meaningful variation, especially when triggered by ambiguity in the premise or hypothesis. While underspecified guidelines and annotator behavior contribute to variation, content-based ambiguity provides a process-independent signal of divergent human perspectives. We call for a shift toward ambiguity-aware NLI that first identifies ambiguous input pairs, classifies their types, and only then proceeds to inference. To support this shift, we present a framework that incorporates ambiguity detection and classification prior to inference. We also introduce a unified taxonomy that synthesizes existing taxonomies, illustrates key subtypes with examples, and motivates targeted detection methods that better align models with human interpretation. Although current resources lack datasets explicitly annotated for ambiguity and subtypes, this gap presents an opportunity: by developing new annotated resources and exploring unsupervised approaches to ambiguity detection, we enable more robust, explainable, and human-aligned NLI systems.

自然语言推理模糊性检测可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。