让视频生成模型直接从噪声状态判断优劣,省去解码开销
Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models

- 用轻量查询聚合头从噪声潜空间提取偏好信号
- 实现最优的偏好判断准确率并支持早期筛选候选
- 适合追求高效高质量视频生成的研究与开发者
当前视频生成评估依赖像素级奖励模型,脱离了噪声扩散过程且成本高昂。本文提出PRISM(Preference Representation in Intermediate States of Diffusion Models),通过冻结的视频扩散主干网络搭配轻量查询聚合头,直接从噪声潜空间解码偏好信号。实验表明,PRISM不仅达到最先进的偏好判断准确率,还具备强噪声鲁棒性,支持在去噪初期即进行Best-of-N采样,提前过滤低质候选,大幅降低计算开销的同时提升视频质量。此外,我们发现主干模型生成能力与其内在评估性能呈强正相关,为自提升视频生成模型提供可能。
原文摘要 · Abstract (English)
Evaluating video generation with clean, pixel-based reward models disconnects evaluation from the noisy diffusion process and incurs massive VAE decoding costs. In this paper, we challenge this paradigm by asking a fundamental question: Can a powerful video generator inherently discriminate preferences directly from noisy latents? To answer this, we introduce \textbf{PRISM} (\textbf{P}reference \textbf{R}epresentation in \textbf{I}ntermediate \textbf{S}tates of Diffusion \textbf{M}odels). PRISM employs a lightweight Query-based Aggregation head with a frozen video diffusion backbone to decode preference signals from noisy latents. Surprisingly, PRISM not only achieves SOTA preference accuracy but also unlocks strong noise-robustness, which enables early-stage Best-of-$N$ sampling. This allows for filtering suboptimal candidates at the very beginning of denoising, drastically reducing computation while boosting video quality. We also reveal a strong positive correlation between a backbone's generative performance and its inherent evaluative power, enabling self-improving video backbones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。