让大模型在多个正确答案时仍能给出引用来源,提升问答可信度。
Adaptive Question Answering: Enhancing Language Model Proficiency for Addressing Knowledge Conflicts with Source Citations
- 提出在模糊答案场景下生成引用的新型问答任务
- 构建5个带引用元数据的新数据集,含多跳推理真实语境
- 提供新评估指标与基线,推动可解释性问答研究
解决知识冲突是问答任务的关键挑战,因互联网存在大量矛盾事实与观点。现有研究或处理多答案模糊场景但忽略引用,或仅针对单一答案场景生成引用,未能兼顾真实复杂性。本文首次提出在多答案模糊设置中进行带引用的问答任务。为此构建了综合性框架:(1) 基于三个阅读理解数据集,通过添加引用元数据生成五个新数据集,涵盖干扰项与改写等模糊情境;(2) 创建首个包含真实自然语境的多跳模糊问答数据集;(3) 设计两个新评估指标;(4) 在五种大语言模型上实现基于规则、提示和微调的多种基线方法。本工作旨在推动社区发展更可信、可解释的问答系统。
原文摘要 · Abstract (English)
Resolving knowledge conflicts is a crucial challenge in Question Answering (QA) tasks, as the internet contains numerous conflicting facts and opinions. While some research has made progress in tackling ambiguous settings where multiple valid answers exist, these approaches often neglect to provide source citations, leaving users to evaluate the factuality of each answer. On the other hand, existing work on citation generation has focused on unambiguous settings with single answers, failing to address the complexity of real-world scenarios. Despite the importance of both aspects, no prior research has combined them, leaving a significant gap in the development of QA systems. In this work, we bridge this gap by proposing the novel task of QA with source citation in ambiguous settings, where multiple valid answers exist. To facilitate research in this area, we create a comprehensive framework consisting of: (1) five novel datasets, obtained by augmenting three existing reading comprehension datasets with citation meta-data across various ambiguous settings, such as distractors and paraphrasing; (2) the first ambiguous multi-hop QA dataset featuring real-world, naturally occurring contexts; (3) two new metrics to evaluate models' performances; and (4) several strong baselines using rule-based, prompting, and finetuning approaches over five large language models. We hope that this new task, datasets, metrics, and baselines will inspire the community to push the boundaries of QA research and develop more trustworthy and interpretable systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。