提出AR-CoPO框架,实现流式自回归视频生成的精准对齐。
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
- 通过分块分支机制构建邻近候选,实现片段级对齐。
- 在Self-Forcing数据集上提升域外泛化与人类偏好对齐效果。
- 适合需要低延迟高质量视频生成与人类反馈对齐的研究者。
流式自回归(AR)视频生成结合少步蒸馏可实现低延迟、高质量合成,但难以通过人类反馈强化学习(RLHF)进行对齐。现有基于SDE的GRPO方法在此场景面临挑战:少步ODE与一致性模型采样器偏离标准流匹配ODE,且其短时、低随机性轨迹对初始化噪声高度敏感,导致中间SDE探索无效。本文提出AR-CoPO(AutoRegressive Contrastive Policy Optimization),将邻居GRPO对比视角适配至流式AR生成。AR-CoPO引入分块对齐机制,在随机选择的片段处构建邻近候选,分配序列级奖励,并执行局部GRPO更新。进一步提出半在线策略训练策略,结合在线策略探索与参考轨迹重放缓存的利用,提升跨领域生成质量。在Self-Forcing上的实验表明,相比基线,AR-CoPO在域外泛化与域内人类偏好对齐方面均有提升,证明了真实对齐而非奖励欺骗。
原文摘要 · Abstract (English)
Streaming autoregressive (AR) video generators combined with few-step distillation achieve low-latency, high-quality synthesis, yet remain difficult to align via reinforcement learning from human feedback (RLHF). Existing SDE-based GRPO methods face challenges in this setting: few-step ODEs and consistency model samplers deviate from standard flow-matching ODEs, and their short, low-stochasticity trajectories are highly sensitive to initialization noise, rendering intermediate SDE exploration ineffective. We propose AR-CoPO (AutoRegressive Contrastive Policy Optimization), a framework that adapts the Neighbor GRPO contrastive perspective to streaming AR generation. AR-CoPO introduces chunk-level alignment via a forking mechanism that constructs neighborhood candidates at a randomly selected chunk, assigns sequence-level rewards, and performs localized GRPO updates. We further propose a semi-on-policy training strategy that complements on-policy exploration with exploitation over a replay buffer of reference rollouts, improving generation quality across domains. Experiments on Self-Forcing demonstrate that AR-CoPO improves both out-of-domain generalization and in-domain human preference alignment over the baseline, providing evidence of genuine alignment rather than reward hacking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。