arXiv:2508.12800cs.CLcs.AI2025-08被引 22

让AI思考更精细,用分步奖励提升复杂问题解决能力

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward

  • 将推理拆解为细粒度思维单元,逐步给予奖励指导
  • 在7个基准上超越现有方法,推理路径更高效可解释
  • 适合需要深度自主研究的AI系统开发者

大型语言模型(LLMs)具备强大问题求解能力,但受限于静态内部知识。检索增强生成(RAG)虽扩展外部信息访问,却在多跳推理和策略搜索方面受限于僵化流程。近期基于智能体的深度研究使LLM能自主推理、搜索与整合信息。然而,依赖结果反馈的强化学习方法面临梯度冲突与奖励稀疏性问题,制约性能提升与训练效率。为此,本文提出原子思维(Atomic Thought)新范式,将推理分解为细粒度功能单元,并由推理奖励模型(RRM)提供原子思维奖励(ATR),实现精准引导。在此基础上,提出Atom-Searcher框架,融合原子思维与ATR,采用类课程式奖励调度,早期优先过程级ATR,后期过渡至结果奖励,加速有效推理路径收敛。在7个基准上的实验表明,该方法持续优于当前最优方案。关键优势包括:(1) 测试时可扩展计算资源;(2) 原子思维为RRM提供监督锚点,连接深度研究任务与奖励模型;(3) 展现更具可解释性的类人推理模式。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit remarkable problem-solving abilities, but struggle with complex tasks due to static internal knowledge. Retrieval-Augmented Generation (RAG) enhances access to external information, yet remains limited in multi-hop reasoning and strategic search due to rigid workflows. Recent advancements in agentic deep research empower LLMs to autonomously reason, search, and synthesize information. However, current approaches relying on outcome-based reinforcement learning (RL) face critical issues such as conflicting gradients and reward sparsity, limiting performance gains and training efficiency. To address these, we first propose Atomic Thought, a novel LLM thinking paradigm that decomposes reasoning into fine-grained functional units. These units are supervised by Reasoning Reward Models (RRMs), which provide Atomic Thought Rewards (ATR) for fine-grained guidance. Building on this, we propose Atom-Searcher, a novel RL framework for agentic deep research that integrates Atomic Thought and ATR. Atom-Searcher uses a curriculum-inspired reward schedule, prioritizing process-level ATR early and transitioning to outcome rewards, accelerating convergence on effective reasoning paths. Experiments on seven benchmarks show consistent improvements over the state-of-the-art. Key advantages include: (1) Atom-Searcher scales computation at test-time. (2) Atomic Thought provides supervision anchors for RRMs, bridging deep research tasks and RRMs. (3) Atom-Searcher exhibits more interpretable, human-like reasoning patterns.

智能体强化学习推理优化LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。