arXiv:2602.07595cs.CVcs.AI2026-02被引 3

让视频生成更精准可控,且稳定可靠。

TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation

  • 分阶段优化,结合监督、强化学习与偏好训练
  • 提升画面质量、时间连贯性与提示遵循度
  • 适合工业级视频生成系统部署

后训练是将预训练视频生成模型转化为可投入生产使用的模型的关键步骤,使其具备指令跟随、可控性及长时程鲁棒性。本文提出一个系统化的后训练框架,将监督策略塑造、基于奖励的强化学习与偏好精炼整合为统一的稳定性约束优化流程。该框架针对视频生成的实际挑战设计,包括高回放成本、时序累积失败模式,以及异质、不确定且常弱区分性的反馈。通过将优化视为分阶段、诊断驱动的过程,而非孤立技巧的堆砌,本文总结出一套提升感知保真度、时间连贯性与提示遵循度的完整方案,同时保持初始设定的可控性。所提出的框架为构建可扩展、稳定、可扩展且在真实场景中有效的后训练流水线提供了清晰蓝图。

原文摘要 · Abstract (English)

Post-training is the decisive step for converting a pretrained video generator into a production-oriented model that is instruction-following, controllable, and robust over long temporal horizons. This report presents a systematical post-training framework that organizes supervised policy shaping, reward-driven reinforcement learning, and preference-based refinement into a single stability-constrained optimization stack. The framework is designed around practical video-generation constraints, including high rollout cost, temporally compounding failure modes, and feedback that is heterogeneous, uncertain, and often weakly discriminative. By treating optimization as a staged, diagnostic-driven process rather than a collection of isolated tricks, the report summarizes a cohesive recipe for improving perceptual fidelity, temporal coherence, and prompt adherence while preserving the controllability established at initialization. The resulting framework provides a clear blueprint for building scalable post-training pipelines that remain stable, extensible, and effective in real-world deployment settings.

视频生成可控生成后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。