用强化学习让大模型更可靠,提高回答可信度。
Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with constraints
- 结合校准置信度与检索搜索,多步反思验证答案
- 在维基数据上训练,提升信心与正确率的一致性
- 适合追求高可信度的开放域问答应用
提升大语言模型的可靠性对实际应用至关重要。本文提出首个将置信度校准与基于检索的搜索相结合的框架——Deliberative Searcher。该智能体在维基百科数据上进行多步反思与验证,并通过强化学习算法,在软可靠性约束下优化准确率。实验表明,该方法显著提升了模型信心与答案正确性之间的一致性,使输出更具可信度。论文将持续更新。
原文摘要 · Abstract (English)
Improving the reliability of large language models (LLMs) is critical for deploying them in real-world scenarios. In this paper, we propose \textbf{Deliberative Searcher}, the first framework to integrate certainty calibration with retrieval-based search for open-domain question answering. The agent performs multi-step reflection and verification over Wikipedia data and is trained with a reinforcement learning algorithm that optimizes for accuracy under a soft reliability constraint. Empirical results show that proposed method improves alignment between model confidence and correctness, leading to more trustworthy outputs. This paper will be continuously updated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。