提出时间比率机制,提升机器人视频模型的组合泛化能力。
Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

- 引入时间比率衡量动作头对未来预测隐状态的依赖程度。
- 在真实任务和LIBERO基准上,显著缓解了分布外组合泛化差距。
- 基于动态注意力调节,仅在规划阶段增强未来条件信号。
生成式视频基础模型具备强大的组合先验,但经过机器人动作数据微调后,世界动作模型(WAMs)和视频动作模型(VAMs)常丧失这些先验,我们称之为视频-动作泛化差距。本文系统评估了VAMs的完整设计空间,发现标准设计无法产生可解释模式。为此,我们提出时间比率(TR),一种基于注意力的度量,反映动作头对未来的隐状态滚动相对于当前帧的依赖强度。TR具有两个关键特性:其一,模型对未来的依赖程度(由TR衡量)能预测其组合泛化能力;其二,该依赖随任务阶段自然变化,在规划阶段关注未来帧,而在精确操作时回归当前帧。基于此,我们提出一种推理时自适应引导方法,利用内在注意力模式,在策略依赖未来滚动时动态增强组合视频条件信号。在LIBERO基准和真实世界任务上验证,该方法有效缓解了分布外-内部(OOD-ID)组合泛化差距。
原文摘要 · Abstract (English)
Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on robotic action data. We refer to this discrepancy as the video-action generalization gap. In this paper, we systematically investigate this gap by evaluating a comprehensive design space of VAMs, demonstrating that standard design choices yield no emergent explanation pattern. To explain this behavior, we introduce the Temporal Ratio (TR), an attention-based measure of how strongly the action head relies on future latent rollouts relative to the anchored current frame. TR has two key properties: first, a model's structural reliance on future-predictive latents, measured via TR, acts as a predictor of its compositional generalization capacity; second, it natively fluctuates based on task phase, shifting attention to future frames during planning and reverting to the present frame for precise manipulation. Finally, based on these findings, we propose an inference-time adaptive guidance method, which exploits this intrinsic feature attention pattern to dynamically amplify compositional video conditioning signals precisely when the policy relies on future rollouts. Evaluated on the LIBERO benchmark and real-world tasks, our approach mitigates the OOD-ID compositional generalization gap. More details: https://umishra.me/temporal-ratio/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。