让机器人动作理解环境影响,提升长序列操作成功率。
EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation

- 将动作指令与环境变化效果结合,学习动态交互语义。
- 在模拟和真实机器人任务中,长程操作成功率显著提升。
- 适合需要理解动作因果关系的复杂操作场景。
有效动作表征对机器人操作至关重要,原始控制轨迹通常噪声大、冗余且难以直接建模。现有方法主要编码动作流自身的结构,将动作在环境中的作用视为隐含。然而,操作的本质是改变世界:同一动作段在不同场景下可能产生不同结果,使动作语义本质上依赖于环境。本文提出EDAR(环境依赖动作表征),将动作标记同时锚定在可执行控制结构和预期视觉变化上。通过将运动命令与其环境条件下的效应耦合,EDAR促使学习到的动作空间捕捉交互语义,而非仅命令级模式。在模拟和真实机器人操作基准上的实验表明,EDAR显著提升了下游策略学习性能,尤其在长时序操作任务中表现突出。结果凸显了将动作表征锚定于可执行控制结构与环境条件化视觉变化的重要性。
原文摘要 · Abstract (English)
Learning effective action representations is critical for robotic manipulation, where raw control trajectories are often noisy, redundant, and difficult to model directly. Existing methods mainly encode the structure of the action stream itself, treating the role of actions in the environment as implicit. Yet manipulation is about changing the world: the same action segment can induce different outcomes under different scene contexts, making action semantics inherently environment-dependent. We propose EDAR, an Environment-Dependent Action Representation that grounds action tokens in both executable control structure and expected visual consequences. By coupling motor commands with their environment-conditioned effects, EDAR encourages the learned action space to capture interaction semantics rather than merely command-level patterns. Experiments on simulated and real-robot manipulation benchmarks demonstrate that EDAR improves downstream policy learning, especially in long-horizon manipulation. These results highlight the importance of grounding action representations in executable control structure and environment-conditioned visual change.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。