arXiv:2608.06994cs.ROcs.AI2026-08

分离动作意图与轨迹,提升机器人任务规划的准确性与可解释性

Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models

论文配图:Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models
图 1 · 摘自论文原文
  • 通过推理链引导显式建模状态变化,解耦高层语义与低层轨迹
  • 在复杂操作任务中成功率显著提升,且支持少样本实机微调
  • 增强物理可解释性,适用于主流世界动作模型迁移

世界动作模型(WAMs)旨在构建统一架构,以理解世界状态演化并指导生成式运动规划。然而现有视觉分支仅预测静态观测,未能捕捉运动交互下的潜在状态演变信息,导致动作模型中高层物理演化与低层轨迹生成存在表征纠缠,形成结构性瓶颈,削弱了对动作生成的世界演化预测能力。本文提出PILOT(用于潜在轨迹优化的物理推理),其核心表征解耦(RD)机制将运动推理链(CoT)作为原生能力集成,促使动作分支显式建模潜在状态转移标记,并将其保留在推理空间中,指导精细运动轨迹生成。实验表明,RD不仅显著提升复杂机器人操作任务中的成功率与泛化能力,还通过解耦高层运动语义与低层轨迹细节,增强了模型的物理可解释性。此外,RD引入的丰富状态转移监督信号有效缓解动作生成中的稀疏监督问题,使其成为高效的少样本实机微调策略,展现出对主流WAM架构的优异可扩展性。

原文摘要 · Abstract (English)

World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative motion planning. However, existing visual branches focus on predicting static visual observation, rather than reflecting potential transition information that captures the evolution of world states under motion interactions. This leads to representational entanglement between high-level physical condition evolution and low-level action trajectory generation within the Action Model, creating a structural bottleneck while weakening the predictive capability of world evolution modeling for action generation. We propose PILOT (Physical Inference for Latent Optimized Trajectories), whose core Representational Deduction (RD) bridges this gap by integrating motion thought-of-chain (CoT) guidance as a native model capability. Specifically, RD aims to encourage the action branch to explicitly model potential state transition tokens, which are retained as CoT in the reasoning space to guide fine-grained motion trajectory. Experiments demonstrate that RD not only significantly improves the success rate and generalization ability of WAMs in complex robotic manipulation tasks but also enhances the model's physical interpretability by decoupling high-level motion semantics from low-level trajectory details. Furthermore, the abundant state transition supervision signals introduced by RD effectively alleviate the sparse supervision in action generation, enabling it to serve as an efficient few-shot real-robot fine-tuning strategy and demonstrating superior scalability for migration to mainstream WAM architectures.

机器人控制动作建模表征解耦少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。