用视觉图像直接控制机器人,让动作更自然、通用性更强。
iMaC: Translating Actions into Motion and Contact Images for Embodied World Models

- 用图像作为动作输入,替代传统关节角度等低维向量。
- 在多个场景中预测准确率和任务成功率显著优于基线方法。
- 无需手动设计动作空间,适合各种不同机器人的通用控制。
具身世界模型已成为视觉机器人决策与交互环境模拟的关键范式。然而,传统框架依赖低维结构化动作向量(如关节角和末端执行器位姿),存在表达能力有限、跨不同机器人泛化性差、复杂物理交互动态建模不自然等问题。为此,本文提出iMaC(Image as Action Control),一种将原始视觉图像作为具身世界模型原生动作表示的统一控制范式。与传统显式运动学动作编码不同,iMaC将连续视觉操作建模为基于图像的动作标记,天然包含空间运动意图、交互几何约束及细微物理动态。我们构建了双分支具身架构:图像动作编码器将目标导向的视觉图像压缩为紧凑的动作嵌入,动态世界预测器则基于图像动作学习环境状态转移规则,实现高保真未来状态预测与闭环具身控制。在公开具身操作基准和真实机器人场景上进行大量实验,结果表明,iMaC在预测精度、任务成功率和跨场景泛化能力上均显著超越基于向量的动作控制基线。此外,该图像动作设计消除了对人工定义动作空间的依赖,实现了异构具身智能体的灵活与通用控制。本工作为具身世界模型提供了创新的视觉-动作视角,提供了一种简单而有效的可扩展机器人感知与操作范式。
原文摘要 · Abstract (English)
Embodied world models have emerged as a pivotal paradigm for visual robotic decision-making and interactive environment simulation. However, conventional embodied frameworks rely on low-dimensional structured action vectors (e.g., joint angles and end-effector poses), which suffer from limited expressive capacity, poor generalization across diverse embodiments, and unnatural dynamic modeling for complex physical interactions. To address these limitations, this paper proposesiMac (Image as Action Control), a novel unified control paradigm that treats raw visual images as native action representations for embodied world models. Departing from traditional explicit kinematic action encoding, iMac formulates continuous visual manipulation as image-based action tokens, which inherently encapsulate spatial motion intentions, interactive geometric constraints and subtle physical dynamics. We construct a dual-branch embodied architecture consisting of an image-action encoder and a dynamic world predictor: the encoder compresses target-driven visual images into compact action embeddings, while the predictor learns environment transition rules conditioned on image actions to achieve high-fidelity future state prediction and closed-loop embodied control. Extensive experiments are conducted on public embodied manipulation benchmarks and real-world robotic scenarios. The results demonstrate that iMac outperforms vector-based action control baselines in prediction accuracy, task success rate and cross-scene generalization ability. Moreover, our image-action design eliminates the reliance on manually defined action spaces, realizing flexible and universal control for heterogeneous embodied agents. This work provides an innovative visual-action perspective for embodied world models, offering a simple yet effective paradigm for scalable robotic perception and manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。