arXiv:2605.14278cs.CV2026-05被引 8

KVPO通过语义探索实现视频生成对齐,提升长时序连贯性。

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration

论文配图:KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
图 1 · 摘自论文原文
  • 用历史键值缓存替代噪声进行语义级多样性探索
  • 在流匹配速度空间中构建奖励加权对比目标,提升一致性
  • 适合追求长视频生成质量与对齐效果的研究者

对齐流式自回归(AR)视频生成器与人类偏好面临挑战。现有强化学习方法多依赖噪声探索和与确定性常微分方程(ODE)动态不匹配的随机微分方程(SDE)代理策略,常扰动低层外观而非关键的高层语义叙事进展,影响长时序一致性。为此,我们提出KVPO,一种面向流式视频生成的原生ODE在线组相对策略优化框架。为实现多样性探索,KVPO引入因果-语义探索范式,将变化源从随机噪声转移到历史键值(KV)缓存,通过随机路由历史KV条目构建语义多样生成分支,且严格保持在数据流形上。在策略建模方面,提出基于轨迹速度能量(TVE)的速度场代理策略,量化分支在流匹配速度空间中的概率,生成与原生ODE形式完全一致的奖励加权对比目标。在多个蒸馏自回归视频生成器上的实验表明,无论单提示短视频还是多提示长视频场景,均在视觉质量、运动质量和文本-视频对齐上取得一致提升。

原文摘要 · Abstract (English)

Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rely on noise-based exploration and SDE-based surrogate policies that are mismatched to the deterministic ODE dynamics of distilled AR models, and tend to perturb low-level appearance rather than the high-level semantic storyline progression critical for long-horizon coherence. To address these limitations, we present KVPO, an ODE-native online Group Relative Policy Optimization (GRPO) framework for aligning streaming video generators. For diversity exploration, KVPO introduces a causal-semantic exploration paradigm that relocates the source of variation from stochastic noise to the historical KV cache. By stochastically routing historical KV entries, it constructs semantically diverse generation branches that remain strictly on the data manifold. For policy modeling, KVPO introduces a velocity-field surrogate policy based on Trajectory Velocity Energy (TVE), which quantifies branch likelihood in flow-matching velocity space and yields a reward-weighted contrastive objective fully consistent with the native ODE formulation. Experiments on multiple distilled AR video generators demonstrate consistent gains in visual quality, motion quality, and text-video alignment across both single-prompt short-video and multi-prompt long-video settings.

视频生成强化学习扩散模型语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。