用单次推理实现高精度动作预测,提升机器人实时操作能力
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
- 通过自蒸馏将多步生成先验压缩至单步推理
- 在仿真与真实环境均超越现有方法,支持高效精准操作
- 适合需要快速决策的机器人操控场景
视频动作模型(VAMs)因其强大的视觉前瞻性,成为机器人学习的有力范式。然而,现有VAMs通常依赖缓慢的多步视频生成或噪声较大的单步特征提取,难以兼顾实时推理与高保真前瞻性。为此,我们提出S-VAM,一种通过单次前向传播预见一致几何与语义表征的快捷视频动作模型。这些预见表征作为稳定蓝图,显著简化动作预测。为实现这一高效捷径,我们引入新颖的自蒸馏策略,将扩散模型多步去噪的结构化生成先验浓缩至单步推理。具体而言,由扩散模型自身生成视频中提取的视觉基础模型(VFM)表征作为教师目标,轻量解耦器作为学生,学习直接将噪声单步特征映射至这些目标。大量仿真与真实世界实验表明,S-VAM优于当前最先进方法,在复杂环境中实现高效且精确的操纵。
原文摘要 · Abstract (English)
Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs, typically relying on either slow multi-step video generation or noisy one-step feature extraction, cannot simultaneously guarantee real-time inference and high-fidelity foresight. To address this limitation, we propose S-VAM, a shortcut video-action model that foresees coherent geometric and semantic representations via a single forward pass. Serving as a stable blueprint, these foreseen representations significantly simplify the action prediction. To enable this efficient shortcut, we introduce a novel self-distillation strategy that condenses structured generative priors of multi-step denoising into one-step inference. Specifically, vision foundation model (VFM) representations extracted from the diffusion model's own multi-step generated videos provide teacher targets. Lightweight decouplers, as students, learn to directly map noisy one-step features to these targets. Extensive experiments in simulation and the real world demonstrate that our S-VAM outperforms state-of-the-art methods, enabling efficient and precise manipulation in complex environments. Our project page is https://haodong-yan.github.io/S-VAM/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。