arXiv:2607.24159cs.ROcs.CV2026-07被引 2

提出解耦视频与动作的机器人策略模型,提升物理动态理解能力。

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

论文配图:DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
图 1 · 摘自论文原文
  • 视频与动作分支独立建模,通过多层特征传递增强信息交流。
  • 引入物理显著性引导(可达性/深度),在少数据下实现更快收敛。
  • 适合需要精准物理推理的机器人操控任务,尤其适用于真实场景部署。

通用机器人操作需要能够预测视觉场景演变并执行语言指令的策略。尽管近期视觉-语言-动作模型通过大规模预训练获得优势,但其主要静态预训练目标对物理动态和时序因果关系监督有限,控制相关知识仍需从下游机器人演示中学习。视频生成模型通过未来预测编码丰富的时空先验,具有潜力。然而现有视频-动作模型或在共享主干中耦合视频与动作预测,使策略适配优化困难;或在引导动作分支时未充分使用视频信息。本文提出DeVA,一种解耦的视频-动作模型,包含专用的视频与动作专家、多层级特征传递及物理显著性引导。DeVA将多个视频层的表征传递至动作专家,实现丰富信息交换的同时使策略学习更可优化,并通过物理显著性引导(可达性/深度)监督中间视频特征与动作流。在仿真基准与真实世界部署实验中均表现优异,仅用少量数据即可快速收敛,且物理引导带来明显性能提升。

原文摘要 · Abstract (English)

Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations. Video generative models offer a promising foundation by encoding rich spatiotemporal priors through future predictions. However, existing Video-Action Models either couple video and action prediction in a shared backbone, making policy adaptation harder to optimize, or under-utilize video information when guiding the action branch. In this work, we introduce DeVA, a Decoupled Video-Action model with specialized video and action experts, multi-level feature transfer, and physically salient guidance. DeVA transfers representations from multiple video layers to the action expert, enabling rich information exchange while making policy learning more tractable. It further supervises intermediate video features and the action stream with physically salient guidance (affordance/depth). Experiments on both simulation benchmarks and real-world deployment demonstrate strong performance with limited data, faster convergence than a unified architecture, and clear performance gains from physical guidance.

机器人策略视频生成物理引导解耦模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。