arXiv:2606.24448cs.RO2026-06被引 1

用生成视频中的几何轨迹指导视觉模型,提升机器人动作学习效果

Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos

论文配图:Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos
图 1 · 摘自论文原文
  • 从人类操作视频生成机器人轨迹,仅用几何信息监督视觉部分
  • 在真实机器人任务上性能超越伪动作基线,接近大量真实数据训练效果
  • 适合研究生成数据与真实控制对齐、视觉-语言-动作模型优化的学者

视觉-语言-动作(VLA)模型需要大规模视频-动作配对数据,但真实远程操控数据稀缺。虽然生成的机器人视频可作为可扩展替代方案,现有方法将合成像素视为真实数据,通过恢复伪动作来训练。我们指出:从生成图像中推导底层控制是一种不匹配的抽象。视频仅保留几何信息(任务的‘位置’),而真实示范包含精确控制信号(任务的‘方式’)。人类到机器人的视频生成过程不对称地保留这些信息:几何信息得以保留,控制信号则丢失。基于此‘非对称保留原理’,我们提出几何引导表示对齐(GRA):通过姿态估计、重定向、仿真和校准投影,从源人类视频提取未来2D末端执行器目标点,并通过辅助2D头注入VLA视觉主干。动作头仅在真实示范上训练。微调时,目标点损失作为空间表征锚点,防止主干丢失几何基础。在真实机器人任务上,GRA在相同数据预算下优于伪动作基线,并缩小了与使用更多真实数据训练策略的差距,表明正确路由的几何信息比恢复的动作更可靠地连接生成视频与机器人策略。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models require large-scale video-action pairs, yet real teleoperation remains scarce. While generated robot videos offer a scalable alternative, existing methods treat them as real robot data by recovering pseudo-actions from synthesized pixels. We argue that deriving low-level control from generated visuals is a mismatched abstraction. A video captures only \emph{geometry}: the spatial trajectory representing the \emph{where} of a task. A real demonstration captures \emph{control}: the exact motor commands representing the \emph{how}. Human-to-robot video generation preserves these unequally: the visible geometry survives the generation process, while the underlying control signals are lost. This \textbf{Asymmetric Preservation Principle} dictates a clean rule: this surviving geometry should solely supervise visual perception, leaving control to real demonstrations. Following this principle, we propose \textbf{GRA} (\textbf{G}eometry-guided \textbf{R}epresentation \textbf{A}lignment), which extracts the geometric content as future 2D end-effector waypoints, computed from the source human video through pose estimation, retargeting, simulation, and calibrated projection, and routes them to the VLA vision backbone via an auxiliary 2D head. The action head is trained on real demonstrations only. During fine-tuning, the waypoint loss persists as a \textbf{spatial representation anchor} that prevents the backbone from losing its geometric grounding. On real-robot tasks, GRA outperforms pseudo-action baselines under matched data budgets and narrows the gap to policies trained with substantially more real demonstrations, suggesting that correctly routed geometry bridges generated videos to robot policies more reliably than recovered actions.

视觉-语言-动作生成数据机器人学习几何引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。