arXiv:2410.16843cs.CL2024-10ICML被引 4

用强化学习让大模型只依赖外部证据,杜绝幻觉。

Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement Learning

  • 通过强化学习训练模型仅依据检索到的外部信息生成回答。
  • 无需人工标注,模型可自主达到可信响应状态。
  • 适合需要高可信度的问答、医疗等关键场景。

可信性是大语言模型在现实应用中的必要前提。本文聚焦于检索增强型语言模型的可信性问题。尽管依赖外部证据,检索增强生成仍存在幻觉,主要源于上下文知识与参数化知识之间的冲突。我们认为,检索增强型语言模型本身具备根据上下文和参数知识生成回应的能力。受对齐人类偏好的启发,我们首次尝试将检索增强型语言模型对齐至仅依赖外部证据、忽略参数知识干扰的状态。具体地,提出基于强化学习的Trustworthy-Alignment算法,从理论和实验上证明大语言模型可在无显式监督下实现可信状态。本工作凸显了大语言模型自我探索内在能力的潜力,并将对齐的应用从满足人类偏好扩展至构建可信智能体。

原文摘要 · Abstract (English)

Trustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect to retrieval augmentation. Despite being supported with external evidence, retrieval-augmented generation still suffers from hallucinations, one primary cause of which is the conflict between contextual and parametric knowledge. We deem that retrieval-augmented language models have the inherent capabilities of supplying response according to both contextual and parametric knowledge. Inspired by aligning language models with human preference, we take the first step towards aligning retrieval-augmented language models to a status where it responds relying merely on the external evidence and disregards the interference of parametric knowledge. Specifically, we propose a reinforcement learning based algorithm Trustworthy-Alignment, theoretically and experimentally demonstrating large language models' capability of reaching a trustworthy status without explicit supervision on how to respond. Our work highlights the potential of large language models on exploring its intrinsic abilities by its own and expands the application scenarios of alignment from fulfilling human preference to creating trustworthy agents.

可信生成强化学习检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。