arXiv:2605.30484cs.RO2026-05

让机器人模型提前预判动作轨迹,提升复杂环境下的操作鲁棒性。

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation

论文配图:ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation
图 1 · 摘自论文原文
  • 用自身关节数据生成未来关键点轨迹,作为时空监督信号
  • 在多个数据集上显著超越基线,尤其在场景变化时表现更优
  • 可插拔设计,不改变原有模型结构,适合部署到现有系统

视觉-语言-动作(VLA)模型在机器人操作中展现潜力,但多数策略仅基于当前观测直接回归动作,未显式建模未来动态,限制了其在分布外扰动下的泛化能力。为此,我们提出ELAN4D,一种以本体为中心、具备4D感知的训练框架,通过引入机器人关键点的未来轨迹作为预测性时空监督,增强VLA策略。仅利用本体感知状态的前向运动学,即可低成本生成关节与末端执行器的3D位移轨迹,提供度量且紧凑的监督信号,无需外部追踪或重建。一个轻量级可插拔的辅助分支搭配轻量解码器,在训练中注入4D信号,通过梯度隔离保留预训练的视觉-语言主干。推理时解码器被移除,基础策略接口保持不变。在LIBERO、LIBERO-Plus、RoboTwin2.0及真实世界任务上的大量实验表明,ELAN4D持续优于强基准模型,在相机、背景和布局变化等分布外扰动下取得显著提升,验证了本体中心4D监督对构建更鲁棒、泛化性更强操作策略的有效性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing actions from current observations, without explicitly modeling future dynamics. This limits their ability to generalize under out-of-distribution perturbations. To address this issue, we propose ELAN4D, an embodiment-centric, 4D-aware training framework that enhances VLA policies with future robot keypoint tracks as predictive spatio-temporal supervision. Using only forward kinematics from proprioceptive states, we derive 3D displacement tracks of robot keypoints, such as joints and the end-effector, with negligible preprocess cost. These tracks provide metric and compact supervision without requiring external trackers or reconstruction. A plug-and-play auxiliary branch with a lightweight track decoder injects this 4D signal into the action expert while preserving the pretrained vision-language backbone through gradient isolation. The track decoder is discarded during inference, leaving the base policy interface unchanged. Extensive experiments on LIBERO, LIBERO-Plus, RoboTwin2.0 and real-world manipulation tasks demonstrate that ELAN4D consistently improves over strong VLA baselines, achieving the best overall performance and substantial gains under out-of-distribution perturbations, including camera, background, and layout shifts. These results highlight the effectiveness of embodiment-centric 4D supervision for building more robust and generalizable manipulation policies.

机器人控制4D监督泛化能力VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。