用强化学习提升问答模型可靠性,减少幻觉。
Enhancing Reliability across Short and Long-Form QA via Reinforcement Learning
- 构建新训练集并设计事实奖励机制,分别应对内外部幻觉。
- 在多个基准上显著降低幻觉率,同时提升长短期问答表现。
- 适合关注大模型可信推理与安全应用的研究者。
尽管强化学习提升了大语言模型的复杂推理能力,但也加剧了幻觉问题,造成能力与可靠性的矛盾。本文提出一种针对性的强化学习框架,旨在缓解短时和长时问答中的内在与外在幻觉。针对外在幻觉(内部知识错误),通过开放生成方式转换TriviaQA构建新型训练集;针对内在幻觉(脱离上下文),利用FineWeb中的长文本设计事实对齐奖励机制。此外,显式奖励模型拒绝回答不可答问题,培养其谨慎性。大量实验表明,该方法在多样化基准上均取得显著性能提升,大幅降低两类幻觉。本研究为解决高级推理与事实可信性之间的核心矛盾提供了实用方案,推动更强大且可靠的大型语言模型发展。
原文摘要 · Abstract (English)
While reinforcement learning has unlocked unprecedented complex reasoning in large language models, it has also amplified their propensity for hallucination, creating a critical trade-off between capability and reliability. This work confronts this challenge by introducing a targeted RL framework designed to mitigate both intrinsic and extrinsic hallucinations across short and long-form question answering. We address extrinsic hallucinations (flawed internal knowledge) by creating a novel training set from open-ended conversions of TriviaQA. Concurrently, we tackle intrinsic hallucinations (unfaithfulness to context) by leveraging long-form texts from FineWeb in a fact-grounding reward scheme. To further bolster reliability, our framework explicitly rewards the model for refusing to answer unanswerable questions, thereby cultivating crucial cautiousness. Extensive experiments demonstrate that our methodology yields significant performance gains across a diverse suite of benchmarks, substantially reducing both hallucination types. Ultimately, this research contributes a practical framework for resolving the critical tension between advanced reasoning and factual trustworthiness, paving the way for more capable and reliable large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。