arXiv:2606.30248cs.CVcs.LG2026-06

用数据流形做奖励模型,免费提升视频生成质量

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation

论文配图:Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
图 1 · 摘自论文原文
  • 通过建模高质量数据的流形结构,生成无需标注的奖励信号
  • 在多个基准上显著减少低级失真,增强细节和运动清晰度
  • 适合关注视频生成质量与计算效率的研究者

近期文本到视频(T2V)扩散模型依赖辅助奖励信号(如奖励模型或DPO)来对齐人类审美并提升真实感。但这些信号带来巨大计算开销,需昂贵的人工标注,且对细粒度局部细节改善有限。本文提出:你的数据流形本质上就是个奖励模型。通过显式建模高质量监督微调(SFT)数据的流形结构,并引导视频隐变量落在该流形上,我们获得密集、可微、近乎零成本的奖励信号,显著提升视频质量,尤其缓解低级失真。模型基于局部坐标编码(LCC),捕捉流形的‘骨架’。但直接应用LCC存在均值回归问题,使隐变量向几何均值偏移,丢失高频细节。为此,我们提出壳层局部坐标编码(Shell-LCC),将流形建模为各向同性的壳层,以匹配真实高密度区域。实验表明,该方法提升真实性,增强高频细节,减少过平滑伪影,缓解运动模糊。

原文摘要 · Abstract (English)

Recent text-to-video (T2V) diffusion models rely heavily on auxiliary reward signals (e.g., via reward models or DPO) to align generated content with human aesthetics and improve realism. These signals, however, incur substantial computational overhead, require costly human annotations, and often yield limited improvement in fine-grained local details. In this paper, we argue that your data manifold is secretly a reward model. By explicitly modeling the manifold structure of high-quality Supervised Fine-Tuning (SFT) data and encouraging video latents to lie on this manifold, we derive dense, differentiable, and nearly cost-free reward signals that significantly improve video quality, particularly in mitigating low-level distortions. Our modeling builds upon Local Coordinate Coding (LCC), which captures the `skeleton' of the manifold. However, directly applying LCC suffers from mean regression, pulling latents toward the geometric mean and losing high-frequency details. We therefore extend it to Shell Local Coordinate Coding (Shell-LCC), which models the manifold `surface' as an isotropic shell to align with the true high-density region. Experiments demonstrate that our approach improves realism, enhances high-frequency details, reduces over-smoothing artifacts, and alleviates motion blur.

文本生成视频流形学习扩散模型奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。