视频模型的隐变量为何能指导动作?关键在时间预测而非像素重建。
What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction
- 用统一探针评估多种编码器,比较其动作相关性。
- 时间预训练模型在视觉保真与动作预测间平衡最优,重建质量高者反而动作可恢复性差。
- 动作感知目标能增强对视觉干扰的鲁棒性,适合机器人任务研究者。
视频世界模型日益用于生成预测性视觉表示,但其隐空间中何种预训练信号带来动作相关结构仍不明确。本文通过统一探针评估涵盖图像自监督、带/不带隐变量预测的视频预训练、基于重构的自编码器、扩散模型及强制动态模型等多种编码器家族。采用共同的逆动力学探针目标,发现动作相关结构主要由时间视频预训练驱动,而非像素重建保真度:具备强像素解码能力的模型可能呈现近乎零的动作可恢复性,而视频自监督编码器在视觉保真与动作预测间始终实现最佳帕累托权衡。对比V-JEPA与VideoMAE表明,主要增益来自自然视频的时间上下文,特征级隐变量预测仅带来较小额外收益。该趋势在机器人基准测试中转移,但CALVIN显示静态环境任务可能因强图像先验掩盖时间结构的重要性。最后,逆动力学监督显著提升对视觉退化的鲁棒性,表明动作感知目标能超越干净场景性能,正则化隐空间几何结构。结果表明,时间预测结构——而非重建保真度——是动作相关视频表征的核心要素。
原文摘要 · Abstract (English)
Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces. We study this question through a unified probe-based evaluation across diverse encoder families, including image-only self-supervision, video pretraining with and without latent prediction, reconstruction-based autoencoders, diffusion models, and shortcut-forcing dynamics models. Using a common inverse-dynamics probing objective, we find that action-relevant structure is driven primarily by temporal video pretraining rather than pixel reconstruction fidelity: models with strong pixel decoding quality can exhibit near-zero action recoverability, while video-pretrained self-supervised encoders consistently achieve the best Pareto trade-off between visual fidelity and action prediction. Comparing V-JEPA and VideoMAE further shows that most gains arise from natural-video temporal context, with feature-level latent prediction providing a smaller additional benefit. These trends transfer across robotic benchmarks, though CALVIN reveals that static-environment tasks can partially mask the importance of temporal structure by allowing strong image priors to suffice. Finally, inverse-dynamics supervision substantially improves robustness to visual corruption, suggesting that action-aware objectives regularize latent geometry beyond clean-setting performance. Our results identify temporal predictive structure -- not reconstruction fidelity -- as the primary ingredient underlying action-relevant video representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。