arXiv:2508.12355cs.CL2025-08被引 5

评测大模型在存在矛盾答案时的识别能力,发现其应对冲突漏洞百出。

Consensus or Conflict? Fine-Grained Evaluation of Conflicting Answers in Question-Answering

  • 用事实核查数据构建真实冲突问答集,标注每对答案的冲突类型。
  • 8个主流大模型在冲突识别上表现差,错误率超50%。
  • 适合研究模型鲁棒性、可信AI或评测系统可靠性的人参考。

大型语言模型在问答任务中表现出色,但多答案问答(MAQA)因可能包含矛盾答案而仍具挑战性。传统问答假设证据一致,而MAQA需处理冲突。现有数据集多依赖合成数据、仅限是非题或自动化标注,缺乏真实性。为此,本文提出新方法,利用事实核查数据构建自然冲突问答基准NATCONFQA,支持识别所有有效答案及具体冲突对。评估8个先进大模型发现,它们在各类冲突场景下均表现脆弱,且采用错误策略应对。该基准为真实冲突下的模型评估提供可靠支撑。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated strong performance in question answering (QA) tasks. However, Multi-Answer Question Answering (MAQA), where a question may have several valid answers, remains challenging. Traditional QA settings often assume consistency across evidences, but MAQA can involve conflicting answers. Constructing datasets that reflect such conflicts is costly and labor-intensive, while existing benchmarks often rely on synthetic data, restrict the task to yes/no questions, or apply unverified automated annotation. To advance research in this area, we extend the conflict-aware MAQA setting to require models not only to identify all valid answers, but also to detect specific conflicting answer pairs, if any. To support this task, we introduce a novel cost-effective methodology for leveraging fact-checking datasets to construct NATCONFQA, a new benchmark for realistic, conflict-aware MAQA, enriched with detailed conflict labels, for all answer pairs. We evaluate eight high-end LLMs on NATCONFQA, revealing their fragility in handling various types of conflicts and the flawed strategies they employ to resolve them.

大模型评测冲突检测问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。