arXiv:2603.13506cs.CV2026-03被引 2

通过平衡模型先验与主体生成能力,用少量数据实现高质量视频生成。

LibraGen: Playing a Balance Game in Subject-Driven Video Generation

  • 以数据质量为支点,结合自动与人工筛选提升训练数据
  • 采用多阶段微调与偏好优化,实现运动连贯性与提示对齐的平衡
  • 适合关注高效、高质量视频生成的研究者与开发者

随着视频生成基础模型(VGFMs)的发展,个性化生成尤其是主体到视频(S2V)生成受到广泛关注。然而,如何在保持VGFMs固有的运动连贯性、视觉美感和提示对齐等先验能力的同时,有效扩展其S2V能力,仍是一大挑战。现有方法常以牺牲某一方面为代价强化另一方。为此,我们提出LibraGen,将S2V能力扩展视为内在能力与新能力之间的平衡游戏。受‘抬高支点,调至平衡’理念启发,我们以数据质量为支点,倡导质量优先策略。构建融合自动与人工过滤的混合数据管道以提升整体数据质量。为进一步调和原生能力与新能力,引入‘调至平衡’后训练范式:在监督微调中同时使用跨对与同对数据,并采用模型合并实现有效权衡。随后设计两种定制化直接偏好优化(DPO)流水线——Consis-DPO与Real-Fake DPO,并进行融合以巩固平衡。推理阶段引入时变动态无分类器引导机制,实现灵活精细控制。实验表明,仅使用千规模训练数据,LibraGen即超越多个开源与商用S2V模型。

原文摘要 · Abstract (English)

With the advancement of video generation foundation models (VGFMs), customized generation, particularly subject-to-video (S2V), has attracted growing attention. However, a key challenge lies in balancing the intrinsic priors of a VGFM, such as motion coherence, visual aesthetics, and prompt alignment, with its newly derived S2V capability. Existing methods often neglect this balance by enhancing one aspect at the expense of others. To address this, we propose LibraGen, a novel framework that views extending foundation models for S2V generation as a balance game between intrinsic VGFM strengths and S2V capability. Specifically, guided by the core philosophy of "Raising the Fulcrum, Tuning to Balance," we identify data quality as the fulcrum and advocate a quality-over-quantity approach. We construct a hybrid pipeline that combines automated and manual data filtering to improve overall data quality. To further harmonize the VGFM's native capabilities with its S2V extension, we introduce a Tune-to-Balance post-training paradigm. During supervised fine-tuning, both cross-pair and in-pair data are incorporated, and model merging is employed to achieve an effective trade-off. Subsequently, two tailored direct preference optimization (DPO) pipelines, namely Consis-DPO and Real-Fake DPO, are designed and merged to consolidate this balance. During inference, we introduce a time-dependent dynamic classifier-free guidance scheme to enable flexible and fine-grained control. Experimental results demonstrate that LibraGen outperforms both open-source and commercial S2V models using only thousand-scale training data.

视频生成S2V平衡优化微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。