arXiv:2605.30350cs.ROcs.LG2026-05被引 1

让机器人视觉理解动作带来的变化,提升抓取泛化能力。

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

论文配图:DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
图 1 · 摘自论文原文
  • 用图像、语言和3D运动流三模态联合训练视觉编码器。
  • 在仿真与真实场景中,下游任务性能最高提升22.5%。
  • 适合需要强动作感知的机器人抓取与视觉-语言-动作模型研究者。

机器人操作依赖于保留动作相关性的感知能力。现有学习框架多基于静态识别或视觉-语言对齐预训练的视觉编码器,将运动理解留到下游策略中。我们提出DynaFLIP,一种将运动理解前置到感知阶段的动力学感知多模态预训练框架。从异构的人类与机器人视频中构建图像-语言-3D运动流三元组,作为训练时监督信号,引导仅图像输入的编码器学习。核心思想是促使三模态在共享超球空间中构成的小单纯形体积——体积越小,对齐越强。为避免几何模糊与平凡坍缩,结合单纯形体积最小化、余弦正则化与对比学习目标。分析表明,DynaFLIP聚焦于操控关键区域。所生成的动力学感知表示可作为通用视觉骨干,在多种下游策略(包括视觉-语言-动作模型)中持续优于基线。在多样化仿真与真实场景中验证,分布外情形下性能最高提升+22.5%。结果表明,当视觉表征不仅捕捉存在物,还编码动作下的世界变化时,机器人泛化能力显著增强。

原文摘要 · Abstract (English)

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image-language-3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space -- a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume minimization with a cosine regularizer and a contrastive objective. Our analyses show that DynaFLIP focuses on control-relevant regions critical for manipulation. The resulting dynamics-aware representations serve as reusable visual backbones and consistently outperform baselines across diverse downstream policies, including VLAs. We validate this across diverse simulation and real-world setups, with gains reaching +22.5% under out-of-distribution scenarios. Our results suggest that robot generalization improves when visual representations are trained to encode not just what is present, but how the world changes under action.

机器人感知多模态学习动力学建模视觉表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。