用自动反馈提升视频生成质量,4小时显卡训练即见效
GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning
- 通过提示驱动构建弱点样本,自动优化训练数据
- 利用视觉语言模型反馈加权样本,平均提升4%表现
- 无需人工标注,适合快速迭代视频生成模型
扩散模型虽显著提升了视频生成质量,但仍需微调以改善实例保持、运动合理性、构图与物理合理性等维度。现有方法依赖人工标注和大量算力,实用性受限。本文提出GigaVideo-1,一种无需额外人工监督的高效微调框架。不引入外部高质量数据,而是通过自动反馈释放预训练视频扩散模型的潜在能力。重点优化数据与优化过程:设计提示驱动的数据引擎,生成多样化的弱点导向训练样本;引入奖励引导训练策略,基于预训练视觉语言模型的反馈自适应加权样本,并加入真实性约束。在VBench-2.0基准上以Wan2.1为基线,覆盖17个评估维度,实验显示仅用4 GPU小时即实现几乎所有维度的持续提升,平均增益约4%。无需人工标注且极少真实数据,证明了方法的有效性与高效性。代码、模型与数据将公开。
原文摘要 · Abstract (English)
Recent progress in diffusion models has greatly enhanced video generation quality, yet these models still require fine-tuning to improve specific dimensions like instance preservation, motion rationality, composition, and physical plausibility. Existing fine-tuning approaches often rely on human annotations and large-scale computational resources, limiting their practicality. In this work, we propose GigaVideo-1, an efficient fine-tuning framework that advances video generation without additional human supervision. Rather than injecting large volumes of high-quality data from external sources, GigaVideo-1 unlocks the latent potential of pre-trained video diffusion models through automatic feedback. Specifically, we focus on two key aspects of the fine-tuning process: data and optimization. To improve fine-tuning data, we design a prompt-driven data engine that constructs diverse, weakness-oriented training samples. On the optimization side, we introduce a reward-guided training strategy, which adaptively weights samples using feedback from pre-trained vision-language models with a realism constraint. We evaluate GigaVideo-1 on the VBench-2.0 benchmark using Wan2.1 as the baseline across 17 evaluation dimensions. Experiments show that GigaVideo-1 consistently improves performance on almost all the dimensions with an average gain of about 4% using only 4 GPU-hours. Requiring no manual annotations and minimal real data, GigaVideo-1 demonstrates both effectiveness and efficiency. Code, model, and data will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。