arXiv:2506.19466cs.AI2025-06

用强化学习提升大模型推理能力,解决检索漂移和冗余问题。

KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models

  • 引入强化学习驱动的多机制协同框架,优化推理流程。
  • 在四个基准上准确率和评分显著提升,最高改善达12.3%。
  • 适合需要复杂推理的问答系统研发与部署人员。

本文提出KunLunBaizeRAG,一种基于强化学习的推理框架,旨在增强大语言模型在复杂多跳问答任务中的推理能力。该框架克服了传统RAG存在的检索漂移、信息冗余和策略僵化等关键缺陷。核心创新包括:由RAG驱动的推理对齐(RDRA)机制、搜索-思考迭代增强(STIE)机制、网络本地智能路由(NLR)机制,以及渐进式混合训练策略。实验结果表明,在四个基准测试中,该框架在精确匹配(EM)和大模型评分(LJ)上均取得显著提升,充分验证了其在复杂推理场景下的鲁棒性与有效性。

原文摘要 · Abstract (English)

This paper introduces KunLunBaizeRAG, a reinforcement learning-driven reasoning framework designed to enhance the reasoning capabilities of large language models (LLMs) in complex multi-hop question-answering tasks. The framework addresses key limitations of traditional RAG, such as retrieval drift, information redundancy, and strategy rigidity. Key innovations include the RAG-driven Reasoning Alignment (RDRA) mechanism, the Search-Think Iterative Enhancement (STIE) mechanism, the Network-Local Intelligent Routing (NLR) mechanism, and a progressive hybrid training strategy. Experimental results demonstrate significant improvements in exact match (EM) and LLM-judged score (LJ) across four benchmarks, highlighting the framework's robustness and effectiveness in complex reasoning scenarios.

大模型推理强化学习RAG多跳问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。