让大模型像侦探一样动态搜索信息,提升复杂问题推理能力
Scent of Knowledge: Optimizing Search-Enhanced Reasoning with Information Foraging
- 用强化学习构建动态检索框架,鼓励模型逐步获取信息
- 在多跳问答和实时网页问答中准确率显著超越基线方法
- 适合需要持续探索、灵活决策的智能推理场景
将外部检索引入大语言模型以克服知识截止问题已成为标准做法。然而,传统检索增强生成采用静态预推理检索策略,难以应对模糊、多步或动态变化的信息需求。近期测试时扩展技术显示出巨大潜力,推动向自适应推理时检索转变。受信息觅食理论启发,我们提出InForage,一种将检索增强推理形式化为动态信息搜寻过程的强化学习框架。与现有方法不同,InForage显式奖励中间检索质量,促使模型通过自适应搜索行为迭代获取并整合信息。为支持训练,我们构建了一个由人类引导的数据集,记录复杂真实网络任务中的迭代搜索与推理轨迹。在通用问答、多跳推理任务及新构建的实时网页问答数据集上的广泛评估表明,InForage性能优于基线方法。结果凸显其在构建稳健、自适应、高效推理代理方面的有效性。
原文摘要 · Abstract (English)
Augmenting large language models (LLMs) with external retrieval has become a standard method to address their inherent knowledge cutoff limitations. However, traditional retrieval-augmented generation methods employ static, pre-inference retrieval strategies, making them inadequate for complex tasks involving ambiguous, multi-step, or evolving information needs. Recent advances in test-time scaling techniques have demonstrated significant potential in enabling LLMs to dynamically interact with external tools, motivating the shift toward adaptive inference-time retrieval. Inspired by Information Foraging Theory (IFT), we propose InForage, a reinforcement learning framework that formalizes retrieval-augmented reasoning as a dynamic information-seeking process. Unlike existing approaches, InForage explicitly rewards intermediate retrieval quality, encouraging LLMs to iteratively gather and integrate information through adaptive search behaviors. To facilitate training, we construct a human-guided dataset capturing iterative search and reasoning trajectories for complex, real-world web tasks. Extensive evaluations across general question answering, multi-hop reasoning tasks, and a newly developed real-time web QA dataset demonstrate InForage's superior performance over baseline methods. These results highlight InForage's effectiveness in building robust, adaptive, and efficient reasoning agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。