arXiv:2502.01523cs.CL2025-02EMNLP被引 18

提出新基准,让模型学会识别问答中的隐藏假设。

CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering

  • 用维基片段检索标注问题的多种可能含义。
  • 考虑条件后准确率提升11.75%,显式给出条件再增7.15%。
  • 适合研究模型幻觉与上下文对齐的学者使用。

用户常假设大语言模型具备相同的上下文认知,导致在问答中省略关键信息,产生模糊问题。基于错误假设的回答可能被误认为幻觉。因此,识别隐含假设至关重要。为此,我们提出条件模糊问答(CondAmbigQA)基准,包含2000个模糊问题及基于条件的评估指标。研究首次将“条件”作为显式上下文约束,通过检索式标注:利用维基百科片段识别给定问题的可能解释,并对应标注答案。实验表明,模型在回答前考虑条件可使准确率提升11.75%,若条件显式提供,再增7.15%。结果表明,看似幻觉的现象可能源于问题本身模糊性,而非模型缺陷,证明条件推理在问答中的有效性,为研究人员提供严谨评估工具。

原文摘要 · Abstract (English)

Users often assume that large language models (LLMs) share their cognitive alignment of context and intent, leading them to omit critical information in question-answering (QA) and produce ambiguous queries. Responses based on misaligned assumptions may be perceived as hallucinations. Therefore, identifying possible implicit assumptions is crucial in QA. To address this fundamental challenge, we propose Conditional Ambiguous Question-Answering (CondAmbigQA), a benchmark comprising 2,000 ambiguous queries and condition-aware evaluation metrics. Our study pioneers "conditions" as explicit contextual constraints that resolve ambiguities in QA tasks through retrieval-based annotation, where retrieved Wikipedia fragments help identify possible interpretations for a given query and annotate answers accordingly. Experiments demonstrate that models considering conditions before answering improve answer accuracy by 11.75%, with an additional 7.15% gain when conditions are explicitly provided. These results highlight that apparent hallucinations may stem from inherent query ambiguity rather than model failure, and demonstrate the effectiveness of condition reasoning in QA, providing researchers with tools for rigorous evaluation.

问答系统模型幻觉上下文对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。