arXiv:2605.08115cs.GRcs.CV2026-05

开源视频生成模型Alice v1用新蒸馏法超越闭源模型,速度更快质量更高。

Alice v1: Distillation-Enhanced Video Generation Surpassing Closed-Source Models

论文配图:Alice v1: Distillation-Enhanced Video Generation Surpassing Closed-Source Models
图 1 · 摘自论文原文
  • 用带分数正则的连贯性蒸馏提升生成质量,不牺牲速度。
  • 5秒720p视频仅需4步去噪,速度比教师模型快7倍,评分提升至91.2。
  • 适合研究视频生成、扩散模型与开源创新的开发者和学者。

我们提出Alice v1,一个140亿参数的开源视频生成模型,通过一致性蒸馏结合分数正则化(rCM)实现顶尖质量。不同于传统蒸馏以质量换速度,rCM蒸馏反而超越教师模型性能。其机制包括:(1) 分数正则项作为模式搜索目标,将概率集中在高质量输出,而非覆盖完整教师分布;(2) 采用针对性合成数据流水线与难例挖掘,强化教师处理不一致的失败模式(如物理、手部、面部);(3) 连贯性约束作为隐式正则化,消除对特定噪声样本的“幸运路径”依赖。Alice v1在H100上以4步去噪生成5秒720p 24fps视频,耗时约8秒,相较教师模型50步提速7倍,VBench评分从84.0(Wan2.2)提升至91.2。该表现超越教师模型及闭源系统(Veo3 ~90,Sora2 ~88),在自动化评测中领先,并在人工偏好测试中具竞争力。我们公开所有模型权重、训练代码、合成数据流水线与评估脚本,推动视频生成领域开源研究。

原文摘要 · Abstract (English)

Wepresent Alice v1, a 14-billion parameter open-source video generation model that achieves state-of-the-art quality through consistency distillation with score regularization (rCM). Contrary to conventional distillation-which trades quality for speed-we demonstrate that rCM-based distillation can exceed teacher model quality. We attribute this to three mechanisms: (1) the score regularization term acts as a mode-seeking objective that concentrates probability mass on high-quality outputs rather than covering the full teacher distribution, (2) our targeted synthetic data pipeline with hard example mining provides training signal specifically for failure modes (physics, hands, faces) that the teacher handles inconsistently, and (3) consistency enforcement acts as implicit regularization, eliminating "lucky path" dependence on specific noise samples. Alice v1 generates 5-second 720p videos at 24fps in 4 denoising steps (~8 seconds on H100), a 7x speedup over the 50-step teacher while improving VBench score from 84.0 (Wan2.2) to 91.2. This surpasses both the teacher and closed-source systems including Veo3 (~90) and Sora2 (~88) on automated benchmarks, with competitive results in human preference studies. We release all model weights, training code, synthetic data pipelines, and evaluation scripts to advance open research in video generation.

视频生成扩散模型蒸馏开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。