arXiv:2607.08639cs.ROcs.CV2026-07被引 10

为机器人控制设计专用视频动作模型,实现少样本泛化。

Native Video-Action Pretraining for Generalizable Robot Control

论文配图:Native Video-Action Pretraining for Generalizable Robot Control
图 1 · 摘自论文原文
  • 用语义对齐的视觉动作分词器替代传统重建模型,提升指令理解与动作精度。
  • 采用因果预训练避免双向模型遗忘问题,支持高频率实时推理。
  • 适合需要快速适应新任务的物理机器人系统,尤其擅长复杂操作场景。

视频动作模型为机器人控制提供了新路径,但将专用于数字内容生成的视频生成模型直接复用于物理环境存在根本缺陷。为此,我们提出面向具身智能的视频动作基础模型LingBot-VA 2.0,基于四大核心设计原则:(1)摒弃传统以重构为主的变分自编码器,引入语义-动作对齐的视觉动作分词器,显著提升后续策略学习中的指令遵循能力与动作精度;(2)针对时序动态的严格因果性,采用从零开始的因果预训练范式,避免双向架构在迁移中常见的灾难性遗忘;(3)为满足高频推理需求,采用稀疏MoE骨干网络,在不牺牲效率的前提下扩展模型容量;(4)通过增强的异步推理机制实现闭环控制,可并行预测未来潜在状态,并基于学习到的前向动力学实时重校正每轮推演。真实世界部署验证了LingBot-VA 2.0作为鲁棒基础模型的能力,其在复杂操作任务上展现出优异的少样本泛化性能。

原文摘要 · Abstract (English)

The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed for digital content creation is inherently inadequate for physical environments. To bridge this gap, we present LingBot-VA 2.0, a video-action foundation model built from the ground up for embodiment. Four core design principles showcase its evolution from LingBot-VA. (1) Departing from traditional reconstruction-focused VAEs, we introduce a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, improving instruction following and action precision in subsequent policy learning. (2) Given the strictly causal nature of temporal dynamics, we adopt a causal pretraining paradigm, training from scratch to circumvent the catastrophic forgetting that frequently occurs when adapting bidirectional architectures. (3) To meet the demands of high-frequency inference, our model employs a sparse MoE backbone, expanding model capacity without compromising efficiency. (4) Real-time closed-loop control is realized through an enhanced asynchronous inference scheme, which predicts future latents in parallel with action execution while re-grounding each rollout on the latest observation via learned forward dynamics. Real-world deployment validates LingBot-VA 2.0 as a robust foundation model, as evidenced by its few-shot generalization across complex manipulation tasks.

机器人控制视频动作具身智能少样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。