arXiv:2508.15761cs.CV2025-08被引 50

Waver可一键生成5-10秒高清视频,支持文本/图像转视频。

Waver: Wave Your Way to Lifelike Video Generation

  • 采用混合流DiT架构,统一处理图文视频生成任务。
  • 生成视频达720p原生分辨率,经上采样至1080p,运动更连贯。
  • 在公开榜单中排名前三,超越多数开源模型。

我们提出Waver,一个高性能的统一图像与视频生成基础模型。Waver可直接生成时长5至10秒、原生分辨率为720p的视频,并进一步上采样至1080p。该模型在单一框架内同时支持文生视频(T2V)、图生视频(I2V)和文生图(T2I)生成。我们引入混合流DiT架构以增强模态对齐并加速训练收敛。为确保训练数据质量,构建了全面的数据清洗流程,并基于多模态大模型(MLLM)手动标注并训练视频质量评估模型,用于筛选高质量样本。此外,提供详细的训练与推理方案,以促进高质量视频生成。基于上述贡献,Waver在复杂运动捕捉方面表现卓越,视频合成中的运动幅度与时间一致性均达到领先水平。值得注意的是,截至2025年7月30日10:00 GMT+8,其在Artificial Analysis的T2V与I2V排行榜中均位列前三,持续优于现有开源模型,且达到或超越主流商业解决方案。我们希望本技术报告能帮助社区更高效地训练高质量视频生成模型,推动视频生成技术发展。官方页面:https://github.com/FoundationVision/Waver。

原文摘要 · Abstract (English)

We present Waver, a high-performance foundation model for unified image and video generation. Waver can directly generate videos with durations ranging from 5 to 10 seconds at a native resolution of 720p, which are subsequently upscaled to 1080p. The model simultaneously supports text-to-video (T2V), image-to-video (I2V), and text-to-image (T2I) generation within a single, integrated framework. We introduce a Hybrid Stream DiT architecture to enhance modality alignment and accelerate training convergence. To ensure training data quality, we establish a comprehensive data curation pipeline and manually annotate and train an MLLM-based video quality model to filter for the highest-quality samples. Furthermore, we provide detailed training and inference recipes to facilitate the generation of high-quality videos. Building on these contributions, Waver excels at capturing complex motion, achieving superior motion amplitude and temporal consistency in video synthesis. Notably, it ranks among the Top 3 on both the T2V and I2V leaderboards at Artificial Analysis (data as of 2025-07-30 10:00 GMT+8), consistently outperforming existing open-source models and matching or surpassing state-of-the-art commercial solutions. We hope this technical report will help the community more efficiently train high-quality video generation models and accelerate progress in video generation technologies. Official page: https://github.com/FoundationVision/Waver.

视频生成扩散模型多模态高质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。