arXiv:2606.20310cs.CV2026-06

让视频生成模型直接从噪声状态判断优劣,省去解码开销

Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models

论文配图:Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models
图 1 · 摘自论文原文
  • 用轻量查询聚合头从噪声潜空间提取偏好信号
  • 实现最优的偏好判断准确率并支持早期筛选候选
  • 适合追求高效高质量视频生成的研究与开发者

当前视频生成评估依赖像素级奖励模型,脱离了噪声扩散过程且成本高昂。本文提出PRISM(Preference Representation in Intermediate States of Diffusion Models),通过冻结的视频扩散主干网络搭配轻量查询聚合头,直接从噪声潜空间解码偏好信号。实验表明,PRISM不仅达到最先进的偏好判断准确率,还具备强噪声鲁棒性,支持在去噪初期即进行Best-of-N采样,提前过滤低质候选,大幅降低计算开销的同时提升视频质量。此外,我们发现主干模型生成能力与其内在评估性能呈强正相关,为自提升视频生成模型提供可能。

原文摘要 · Abstract (English)

Evaluating video generation with clean, pixel-based reward models disconnects evaluation from the noisy diffusion process and incurs massive VAE decoding costs. In this paper, we challenge this paradigm by asking a fundamental question: Can a powerful video generator inherently discriminate preferences directly from noisy latents? To answer this, we introduce \textbf{PRISM} (\textbf{P}reference \textbf{R}epresentation in \textbf{I}ntermediate \textbf{S}tates of Diffusion \textbf{M}odels). PRISM employs a lightweight Query-based Aggregation head with a frozen video diffusion backbone to decode preference signals from noisy latents. Surprisingly, PRISM not only achieves SOTA preference accuracy but also unlocks strong noise-robustness, which enables early-stage Best-of-$N$ sampling. This allows for filtering suboptimal candidates at the very beginning of denoising, drastically reducing computation while boosting video quality. We also reveal a strong positive correlation between a backbone's generative performance and its inherent evaluative power, enabling self-improving video backbones.

视频生成扩散模型偏好学习高效采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。