arXiv:2505.16483cs.CLcs.AI2025-05AAAI被引 4

用合成数据和强化学习提升大模型答话忠实度

Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning

  • 用四种任务自动生成可验证的问答数据,无需人工标注
  • 提出双路径强化学习方法,同时优化短/长文本生成忠实度
  • 在11个任务中显著减少幻觉,优于GPT-4o等顶尖模型

让大语言模型在给定上下文中保持忠实对构建可靠信息查询系统至关重要。为此,我们提出系统性框架CANOE,通过无须人工标注的方式,在多个下游任务中降低大模型的忠实性幻觉。具体而言,我们首先利用四种多样化任务自动生成短形式问答(QA)数据,构建高质量且易于验证的训练数据。同时,提出基于规则的双重强化学习方法Dual-GRPO,其包含三个由合成短形式问答数据导出的定制化奖励信号,同步优化短文本与长文本生成。值得注意的是,Dual-GRPO无需人工标注偏好数据来训练奖励模型,且避免了仅依赖合成短形式数据时对短文本生成的过度优化。实验结果表明,CANOE在11个不同任务中显著提升大模型的忠实度,甚至优于最先进的模型如GPT-4o和OpenAI o1。

原文摘要 · Abstract (English)

Teaching large language models (LLMs) to be faithful in the provided context is crucial for building reliable information-seeking systems. Therefore, we propose a systematic framework, CANOE, to reduce faithfulness hallucinations of LLMs across different downstream tasks without human annotations. Specifically, we first synthesize short-form question-answering (QA) data with four diverse tasks to construct high-quality and easily verifiable training data without human annotation. Also, we propose Dual-GRPO, a rule-based reinforcement learning method that includes three tailored rule-based rewards derived from synthesized short-form QA data, while simultaneously optimizing both short-form and long-form response generation. Notably, Dual-GRPO eliminates the need to manually label preference data to train reward models and avoids over-optimizing short-form generation when relying only on the synthesized short-form QA data. Experimental results show that CANOE greatly improves the faithfulness of LLMs across 11 different tasks, even outperforming the most advanced LLMs, e.g., GPT-4o and OpenAI o1.

大模型忠实性强化学习合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。