让机器人通过一次示范视频精准执行任务,关键在预判手部轨迹可见性。
SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos

- 用视觉可见性预测未来手部运动路径,实现空间精确定位。
- 在四个任务设置中成功率最高,真实场景提升12.5个百分点。
- 适用于需要高精度动作的单次示范学习,适合工业或服务机器人。
视觉-语言-动作模型(VLAs)是通用机器人策略的有前途方向,但适应新任务通常需昂贵的任务特定遥控数据。作为替代方案,我们研究了一次示范条件下的VLAs,即机器人策略基于单一未见过任务的示范视频进行条件化。我们发现,现有端到端方法在成功执行需精确定位小目标区域的任务时表现不佳。为解决此问题,我们提出SeeTraceAct,一种示范条件化的VLA框架,通过可视性感知的未来末端执行器轨迹预测,促进精确的空间定位。为支持可复现的跨实体演示评估,我们引入并发布了RoboCasa-DC,即带有成对人形视频的RoboCasa扩展版本。在RoboCasa-DC和一个真实世界基准(使用Franka Panda机械臂,以人类示范为条件)上的实验表明,SeeTraceAct优于基线,在所有四个RoboCasa-DC设置中取得最佳成功率,并在真实世界平均成功率上提升12.5个百分点。
原文摘要 · Abstract (English)
Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-specific teleoperation data. As an alternative, we study one-shot demo-conditioned VLAs, where a robot policy is conditioned on a single demonstration video of an unseen task. We find that existing end-to-end approaches often struggle when successful execution requires precisely localizing small target regions. To address this limitation, we propose SeeTraceAct, a demo-conditioned VLA framework that encourages precise spatial grounding through visibility-aware prediction of future end-effector traces. To enable reproducible evaluation with cross-embodiment demonstrations, we introduce and release RoboCasa-DC, a demo-conditioned extension of RoboCasa with episode-paired humanoid videos. Experiments on RoboCasa-DC and a real-world benchmark, where a Franka Panda arm is conditioned on human demonstrations, show that SeeTraceAct outperforms baselines, achieving the best success rate across all four RoboCasa-DC settings and improving real-world average success by 12.5 percentage points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。