arXiv:2602.04496cs.AI2026-02被引 2

让AI像专家一样思考,通过动态反思和信心控制提升科学推理能力

ReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence Control

  • 采用分阶段的解题-批评-选择架构,根据信心动态调整计算资源
  • 在HLE等基准上超越现有模型,实现专家级推理性能
  • 无需人工标注,自动生成高质量训练数据支持持续优化

大型语言模型在专家级科学推理任务中仍面临挑战,尤其在Humanity's Last Exam(HLE)等基准上,固定的工具流程、脆弱的多智能体协作以及低效的测试时扩展常导致性能受限。我们提出ReThinker,一种具备信心感知的代理框架,通过分阶段的Solver-Critic-Selector架构,动态协调检索、工具使用与多智能体推理。该框架不依赖固定流程,而是基于模型信心动态分配计算,实现自适应工具调用、引导式多维度反思及鲁棒的信心加权选择。为实现无监督可扩展训练,我们进一步提出逆向数据合成管道与自适应轨迹复用策略,将成功推理路径转化为高质量监督信号。在HLE、GAIA和XBench上的实验表明,ReThinker持续优于现有带工具的基础模型与深度研究系统,在专家级推理任务中达到最新水平。

原文摘要 · Abstract (English)

Expert-level scientific reasoning remains challenging for large language models, particularly on benchmarks such as Humanity's Last Exam (HLE), where rigid tool pipelines, brittle multi-agent coordination, and inefficient test-time scaling often limit performance. We introduce ReThinker, a confidence-aware agentic framework that orchestrates retrieval, tool use, and multi-agent reasoning through a stage-wise Solver-Critic-Selector architecture. Rather than following a fixed pipeline, ReThinker dynamically allocates computation based on model confidence, enabling adaptive tool invocation, guided multi-dimensional reflection, and robust confidence-weighted selection. To support scalable training without human annotation, we further propose a reverse data synthesis pipeline and an adaptive trajectory recycling strategy that transform successful reasoning traces into high-quality supervision. Experiments on HLE, GAIA, and XBench demonstrate that ReThinker consistently outperforms state-of-the-art foundation models with tools and existing deep research systems, achieving state-of-the-art results on expert-level reasoning tasks.

科学推理智能体系统信心控制自我反思

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。