让多模态智能体提前规划动作轨迹,提升复杂任务的执行稳定性。
Anticipatory Planning for Multimodal AI Agents
- 分两阶段训练:先预测动作序列,再根据执行反馈微调细节
- 在7个基准上显著提升规划稳定性和任务成功率
- 适合需要长期推理和多步规划的应用场景
多模态智能体在计算机使用和工具操作方面取得进展,但多数系统仍为被动响应,孤立优化动作而缺乏对未来状态或长期目标的推理,导致规划不连贯,难以完成高阶多步任务。本文提出TraceR1,一种两阶段强化学习框架,通过预判短期动作轨迹来显式训练前瞻性推理。第一阶段采用轨迹级强化学习,以奖励确保预测动作序列的全局一致性;第二阶段通过冻结工具代理的执行反馈进行接地强化微调,提升每一步的准确性和可执行性。TraceR1在七个基准上评估,涵盖在线与离线计算机使用、多模态工具推理任务,相比反应式及单阶段基线,在规划稳定性、执行鲁棒性和泛化能力上均有显著提升。结果表明,前瞻性轨迹推理是构建能有效推理、规划并行动的多模态智能体的关键原则。
原文摘要 · Abstract (English)
Recent advances in multimodal agents have improved computer-use interaction and tool-usage, yet most existing systems remain reactive, optimizing actions in isolation without reasoning about future states or long-term goals. This limits planning coherence and prevents agents from reliably solving high-level, multi-step tasks. We introduce TraceR1, a two-stage reinforcement learning framework that explicitly trains anticipatory reasoning by forecasting short-horizon trajectories before execution. The first stage performs trajectory-level reinforcement learning with rewards that enforce global consistency across predicted action sequences. The second stage applies grounded reinforcement fine-tuning, using execution feedback from frozen tool agents to refine step-level accuracy and executability. TraceR1 is evaluated across seven benchmarks, covering online computer-use, offline computer-use benchmarks, and multimodal tool-use reasoning tasks, where it achieves substantial improvements in planning stability, execution robustness, and generalization over reactive and single-stage baselines. These results show that anticipatory trajectory reasoning is a key principle for building multimodal agents that can reason, plan, and act effectively in complex real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。