arXiv:2609.05892cs.RO2026-09

用4D动态轨迹表示交互几何演化,实现人与机器人间操作知识迁移。

A4A: Cross-Embodiment Transfer of Action-Oriented 4D Affordances from Human Demonstrations

论文配图:A4A: Cross-Embodiment Transfer of Action-Oriented 4D Affordances from Human Demonstrations
图 1 · 摘自论文原文
  • 提出动作导向的4D affordance,捕捉交互点随时间演变的轨迹。
  • 在仿真与真实场景中,预训练提升多种视觉-语言-动作模型性能。
  • 适合研究人机协作、机器人模仿学习与跨体感控制的开发者。

人类示范包含丰富的操作知识,但其中哪些信息可有效迁移至机器人控制仍不明确。现有亲和力表征多为2D掩码、3D区域、接触点或可操作性评分,主要标识交互可能发生的位置,却难以建模任务执行过程中交互相关几何的变化。为此,本文提出动作导向的4D亲和力,以语言条件下的交互关键3D点未来轨迹形式表示,捕捉任务相关的几何演化,而非特定本体的动作。该表示支持跨人类与机器人形态的知识迁移。基于此,我们从现有的人-物交互视频数据及互补的RGB-D示范数据构建大规模动作导向4D亲和力数据集,并提出A4A框架——利用4D亲和力轨迹预测进行机器人策略预训练,再进行操作微调。仿真与真实世界实验验证了A4A的有效性,表明使用4D亲和力数据预训练能持续提升多种视觉-语言-动作(VLA)策略的操作性能。结果确立动作导向4D亲和力作为从人类示范向机器人控制迁移操作知识的有效跨体感表示。

原文摘要 · Abstract (English)

Human demonstrations contain rich manipulation knowledge, but it remains unclear what information can be transferred effectively to robot control. Existing affordance representations are typically formulated as 2D masks, 3D regions, contact points, or actionability scores, and therefore primarily identify where interaction may occur. However, effective manipulation also requires modeling how interaction-relevant geometry evolves during task execution. To bridge this gap, we introduce action-oriented 4D affordances, which represent the language-conditioned future trajectories of interaction-relevant 3D points. These trajectories capture task-conditioned geometric evolution rather than embodiment-specific actions, enabling transferable interaction priors across humans and robots. Based on this representation, we construct a large-scale action-oriented 4D affordance dataset from existing human--object interaction video data and complementary RGB-D demonstrations, and introduce A4A, an affordance-to-action framework that uses 4D affordance trajectory prediction to pretrain robot policies before manipulation fine-tuning. Experiments in both simulation and the real world validate the effectiveness of A4A, showing that pretraining with action-oriented 4D affordance data consistently improves the manipulation performance of diverse VLA policies. These results establish action-oriented 4D affordances as an effective cross-embodiment representation for transferring manipulation knowledge from human demonstrations to robot control.

机器人操作4D亲和力知识迁移模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。