用智能验证器提升多模态智能体的推理能力,让其更准确、更可信。
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
- 引入可自适应选择评分函数的验证代理Argos,评估答案、定位和推理过程。
- 在空间推理、机器人等任务上达到最新最佳性能,显著减少幻觉与奖励作弊。
- 适合研究多模态强化学习、智能体推理与具身AI的开发者与研究人员。
以多模态强化学习(MMRL)训练的智能体推理模型日益强大,但普遍依赖基于最终答案的稀疏奖励信号。利用推理过程中的中间标记计算更丰富的奖励,能显著提升学习效果。然而,在MMRL中设计更具信息量的奖励仍具挑战:不同样本需不同评分函数,且教师模型可能提供噪声信号。本文提出Argos(Agentic Reward for Grounded & Objective Scoring),一个原则性的奖励代理,用于训练多模态推理模型。针对每个样本,Argos从一组由教师模型生成和规则驱动的评分函数中选择,同时评估:(i) 最终回答准确性,(ii) 被指代实体与动作的空间时间定位精度,(iii) 推理过程质量。实验表明,结合SFT数据筛选与强化学习训练,使用Argos的模型在空间推理、视觉幻觉及机器人与具身智能基准测试中均达到当前最优表现。关键发现:仅依赖高质量推理数据的SFT后训练不足以防止强化学习阶段出现无根基解;而在线验证机制可有效缓解奖励欺骗。此外,通过帕累托最优理论为Argos的有效性提供了形式化支持。
原文摘要 · Abstract (English)
Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost universally optimized using sparse, outcome-based rewards computed based on the final answers. Richer rewards computed from the reasoning tokens can improve learning significantly by providing more fine-grained guidance. However, it is challenging to compute more informative rewards in MMRL beyond those based on outcomes since different samples may require different scoring functions and teacher models may provide noisy reward signals too. In this paper, we introduce the Argos (Agentic Reward for Grounded & Objective Scoring), a principled reward agent to train multimodal reasoning models for agentic tasks. For each sample, Argos selects from a pool of teacher-model derived and rule-based scoring functions to simultaneously evaluate: (i) final response accuracy, (ii) spatiotemporal localization of referred entities and actions, and (iii) the quality of the reasoning process. We find that by leveraging our agentic verifier across both SFT data curation and RL training, our model achieves state-of-the-art results across multiple agentic tasks such as spatial reasoning, visual hallucination as well as robotics and embodied AI benchmarks. Critically, we demonstrate that just relying on SFT post-training on highly curated reasoning data is insufficient, as agents invariably collapse to ungrounded solutions during RL without our online verification. We also show that our agentic verifier can help to reduce reward-hacking in MMRL. Finally, we also provide a theoretical justification for the effectiveness of Argos through the concept of pareto-optimality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。