arXiv:2507.13942cs.CVcs.AI2025-07被引 4

统一评估冻结视觉模型的预测能力,覆盖从像素到物体运动的多任务场景。

Frozen Forecasting: A Unified Evaluation

  • 在视觉表征空间中训练扩散模型直接预测未来特征
  • 视频预训练模型在各抽象层级上表现优于图像模型
  • 语言监督对预测能力提升不显著,适合研究通用视觉系统

预测未来事件是通用系统在不同抽象层次规划或行动的核心能力。然而,由于未来的固有不确定性,判断预测是否“正确”仍具挑战。本文提出一个统一评估框架,用于衡量冻结视觉主干在多样化任务和抽象层级上的预测能力。不同于关注单个时间步,该框架评估完整轨迹,并引入分布性指标以更好捕捉未来结果的多模态特性。给定一个冻结视觉模型,我们训练潜在扩散模型直接在其表征空间中预测未来特征,再通过轻量级、任务特定的读出层解码。这实现了在一系列多样化任务中的一致评估,同时隔离了主干自身的预测能力。我们将框架应用于九种不同视觉模型,涵盖图像与视频预训练、对比与生成目标,以及有无语言监督,评估其在四个预测任务上的表现,从低层像素预测到高层物体运动。结果发现,预测性能与感知质量强相关;视频合成模型的预测能力在所有抽象层级上均达到或超过掩码预训练模型水平。但语言监督并未持续提升预测效果。值得注意的是,视频预训练模型始终优于图像基模型。

原文摘要 · Abstract (English)

Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, evaluating whether a forecast is "correct" remains challenging due to the inherent uncertainty of the future. We propose a unified evaluation framework for assessing the forecasting capabilities of frozen vision backbones across diverse tasks and abstraction levels. Rather than focusing on single time steps, our framework evaluates entire trajectories and incorporates distributional metrics that better capture the multimodal nature of future outcomes. Given a frozen vision model, we train latent diffusion models to forecast future features directly in its representation space, which are then decoded via lightweight, task-specific readouts. This enables consistent evaluation across a suite of diverse tasks while isolating the forecasting capacity of the backbone itself. We apply our framework to nine diverse vision models, spanning image and video pretraining, contrastive and generative objectives, and with or without language supervision, and evaluate them on four forecasting tasks, from low-level pixel predictions to high-level object motion. We find that forecasting performance strongly correlates with perceptual quality and that the forecasting abilities of video synthesis models are comparable or exceed those pretrained in masking regimes across all levels of abstraction. However, language supervision does not consistently improve forecasting. Notably, video-pretrained models consistently outperform image-based ones.

视觉预测扩散模型冻结主干多任务评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。