提出四阶段后训练框架,提升视频生成的可控性与质量
A Systematic Post-Train Framework for Video Generation

- 分四步优化:指令微调、人类反馈强化学习、提示增强、推理优化
- 在保持低采样成本前提下显著减少伪影,提升画面连贯性与美观度
- 适合需要高可控性视频生成的工业级应用开发者
尽管大规模视频扩散模型在生成高分辨率、语义丰富的内容方面表现出色,但其预训练性能与实际部署需求之间仍存在显著差距,主要问题包括提示敏感、时间不一致和高昂的推理成本。为此,我们提出一个系统的后训练框架,通过四个协同阶段将预训练模型与用户意图对齐:首先采用监督微调(SFT)使基础模型具备稳定的指令遵循能力;接着通过人类反馈强化学习(RLHF)结合专为视频扩散设计的组相对策略优化(GRPO)方法,提升感知质量和时间连贯性;随后利用专用语言模型进行提示增强以优化用户输入;最后通过推理优化提升系统效率。这四个组件共同构建了一个系统化方案,有效提升视觉质量、时间一致性与指令跟随能力,同时保留预训练中学习到的可控性。大量实验表明,该统一流程能有效缓解常见伪影,显著提升可控性与视觉美感,且严格遵守采样成本约束。
原文摘要 · Abstract (English)
While large-scale video diffusion models have demonstrated impressive capabilities in generating high-resolution and semantically rich content, a significant gap remains between their pretraining performance and real-world deployment requirements due to critical issues such as prompt sensitivity, temporal inconsistency, and prohibitive inference costs. To bridge this gap, we propose a comprehensive post-training framework that systematically aligns pretrained models with user intentions through four synergistic stages: we first employ Supervised Fine-Tuning (SFT) to transform the base model into a stable instruction-following policy, followed by a Reinforcement Learning from Human Feedback (RLHF) stage that utilizes a novel Group Relative Policy Optimization (GRPO) method tailored for video diffusion to enhance perceptual quality and temporal coherence; subsequently, we integrate Prompt Enhancement via a specialized language model to refine user inputs, and finally address system efficiency through Inference Optimization. Together, these components provide a systematic approach to improving visual quality, temporal coherence, and instruction following, while preserving the controllability learned during pretraining. The result is a practical blueprint for building scalable post-training pipelines that are stable, adaptable, and effective in real-world deployment. Extensive experiments demonstrate that this unified pipeline effectively mitigates common artifacts and significantly improves controllability and visual aesthetics while adhering to strict sampling cost constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。