用扩散模型统一视频生成与理解,零微调实现多任务高精度感知
Gen4U: Unifying Video Generation and Understanding via Diffusion

- 利用互KNN度量分析扩散模型潜空间,发现语义与细节随噪声层级演化
- 中等噪声下语义线性可分,低噪声下细节分散需注意力机制解码
- 框架仅一次前向传播即适配生成与理解,支持多任务无微调部署
先前研究表明,扩散表示能捕捉低层几何信息但难以处理高层语义。我们证明当前最先进的视频扩散模型克服了这一局限。通过使用最新的互KNN对齐度量系统探测其中间激活,揭示了高度结构化的潜空间:视觉表示随网络深度与噪声水平演变。结果显示,中等噪声水平下全局语义呈线性可分,而细粒度细节在低噪声时仍存在但空间上分散,需注意力机制解码。基于此,我们提出Gen4U(Generation for Understanding)框架,以单次前向传播重用生成式表示。实验表明,冻结的大规模视频扩散模型在多种任务中表现优异,涵盖语义与非语义目标(视频分类、深度估计、相机位姿估计、图像与视频描述生成)。无需微调,Gen4U统一生成与理解范式,在保持高质量视频生成能力的同时,实现强大感知性能。
原文摘要 · Abstract (English)
Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematically probing their intermediate activations using recent mutual-kNN alignment metrics, we reveal a highly structured latent space where visual representations evolve across both network depth and noise levels. We show that while moderate noise levels yield linearly separable global semantics, fine-grained details persist at lower noise levels but become spatially scattered, requiring attention mechanisms to decode. Building on these insights, we introduce Gen4U (Generation for Understanding), a framework that repurposes these generative representations with a single forward pass. Our experiments establish that frozen, large-scale video diffusion models function as highly competitive video encoders across a wide spectrum of tasks, spanning semantic and non-semantic objectives (video classification, depth estimation, camera pose estimation, image and video captioning). Bypassing fine-tuning, Gen4U unifies the generation and understanding paradigms, achieving strong perception performance while fully preserving the model's ability to generate high-quality video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。