测试大模型在带引用的模糊问答中表现,发现其难处理多重正确答案且引用准确率为零。
Factuality or Fiction? Benchmarking Modern LLMs on Ambiguous QA with Citations
- 在模糊问答中引入冲突感知提示,提升模型应对多正确答案能力
- 所有模型引用生成准确率均为0,说明无法正确生成来源引用
- 新提示策略显著改善引用准确性,同时保持正确答案预测能力
评估现代大语言模型在复杂现实任务中的事实准确性与引用性能至关重要。本文针对带有源引用的模糊问答任务,使用三个最新数据集——DisentQA-DupliCite、DisentQA-ParaCite 和 AmbigQA-Cite,分析 GPT-4o-mini 与 Claude-3.5 两款领先模型的表现。结果表明,尽管更大更新的模型在模糊情境中能至少预测一个正确答案,但在面对多个有效答案时仍表现不佳;所有模型在引用生成上均表现极差,引用准确率始终为0。引入冲突感知提示后,模型在处理多重有效答案和提升引用准确率方面实现显著改进,同时保持正确答案预测能力。研究揭示了大模型在处理模糊性与可靠引用方面的挑战与机遇,为构建可信、可解释的问答系统提供了关键洞见与基准。
原文摘要 · Abstract (English)
Benchmarking modern large language models (LLMs) on complex and realistic tasks is critical to advancing their development. In this work, we evaluate the factual accuracy and citation performance of state-of-the-art LLMs on the task of Question Answering (QA) in ambiguous settings with source citations. Using three recently published datasets-DisentQA-DupliCite, DisentQA-ParaCite, and AmbigQA-Cite-featuring a range of real-world ambiguities, we analyze the performance of two leading LLMs, GPT-4o-mini and Claude-3.5. Our results show that larger, recent models consistently predict at least one correct answer in ambiguous contexts but fail to handle cases with multiple valid answers. Additionally, all models perform equally poorly in citation generation, with citation accuracy consistently at 0. However, introducing conflict-aware prompting leads to large improvements, enabling models to better address multiple valid answers and improve citation accuracy, while maintaining their ability to predict correct answers. These findings highlight the challenges and opportunities in developing LLMs that can handle ambiguity and provide reliable source citations. Our benchmarking study provides critical insights and sets a foundation for future improvements in trustworthy and interpretable QA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。