arXiv:2608.24101cs.RO2026-08

用视觉轨迹连接机器人控制与预测,提升任务成功率。

TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks

论文配图:TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks
图 1 · 摘自论文原文
  • 以视觉轨迹为中间桥梁,统一控制与预测信号
  • 仿真和真实场景下成功率分别提升至55%和76%
  • 适合需要精准动作规划的机器人任务

机器人动作具有高度具身性,与图像空间变化关联弱,难以作为世界模型的条件信号。相比之下,视觉轨迹提供了不依赖具身性的运动表征,能提供密集且空间精确的未来视频预测引导。基于此,我们提出TrAct:一种基于世界模型的机器人决策框架,使用视觉轨迹作为控制与预测之间的中间接口。TrAct包含三个组件:视觉-语言-动作-轨迹模型(VLAT),联合预测候选动作与对应视觉轨迹;轨迹条件世界模型(TWM),根据所提轨迹预测未来视觉结果;视觉-语言奖励模型(VLAC),评估预测结果与指令的匹配度。推理时,VLAT生成动作-轨迹对,TWM推演其视觉后果,VLAC选择最符合指令的轨迹,执行对应的行动。在新提出的LIBERO-INTEGRAL基准与真实的Franka操作任务上,相比强基线π_{0.5},TrAct将仿真成功率从27%提升至55%,真实任务从49%提升至76%。此外,TWM在视频预测质量上持续优于动作条件世界模型(AWM)。结果表明,视觉轨迹是连接控制与预测的有效共享接口,可实现更准确的世界建模与更强的机器人泛化能力。

原文摘要 · Abstract (English)

Robot actions are inherently embodiment-specific and only weakly aligned with image-space visual changes, limiting their effectiveness as conditioning signals for robot world models. In contrast, visual tracks provide an embodiment-agnostic representation of how task-relevant points move through a scene, offering dense image-space guidance for accurate and spatially precise future video prediction. Building on this observation, we propose TrAct, a world-model-based robot decision-making framework that uses visual tracks as an intermediate interface between control and prediction. TrAct consists of three components: a Vision-Language-Action-and-Track model (VLAT) that jointly predicts candidate actions and corresponding visual tracks from the current observation and language instruction; a track-conditioned world model (TWM) that predicts future visual outcomes conditioned on the proposed tracks; and a vision-language reward model (VLAC) that scores the predicted outcomes. At inference time, VLAT generates candidate action-track pairs, TWM rolls out their visual consequences, and VLAC selects the track whose predicted outcome best satisfies the instruction; the action paired with the selected track is then executed by the robot. Experiments on the proposed LIBERO-INTEGRAL benchmark and real-world Franka manipulation show that TrAct improves success rates from 27% to 55% in simulation and from 49% to 76% on real-world tasks compared with the strong VLA baseline $π_{0.5}$. Furthermore, TWM consistently improves video prediction quality over the action-conditioned world model (AWM). These results demonstrate that visual tracks provide an effective shared interface between robot control and visual prediction, enabling more accurate world modeling and stronger robot generalization.

机器人控制视觉轨迹世界模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。