Seedance 1.0实现高速高质视频生成,兼顾指令遵循与运动自然性。
Seedance 1.0: Exploring the Boundaries of Video Generation Models
- 融合多源数据与精准描述词,提升模型对复杂场景的学习能力。
- 支持多镜头生成,1080p 5秒视频生成仅需41.4秒,速度提升约10倍。
- 适用于需要高质量、长时序连贯视频的创作者与研究者。
扩散模型的突破推动了视频生成技术的快速发展,但现有基础模型在同时平衡提示遵循性、运动合理性与视觉质量方面仍面临挑战。本文提出Seedance 1.0,一个高性能且推理高效的视频生成基础模型,集成多项核心技术改进:(i) 通过精确且有意义的视频描述增强的多源数据筛选,实现跨多样化场景的全面学习;(ii) 提出高效架构设计与训练范式,原生支持多镜头生成,并联合学习文生视频与图生视频任务;(iii) 采用细粒度监督微调与基于多维度奖励机制的视频专用强化学习人类反馈(RLHF),实现综合性能提升;(iv) 通过多阶段知识蒸馏与系统级优化,实现约10倍的推理加速。Seedance 1.0可在NVIDIA-L20上以41.4秒完成1080p分辨率5秒视频生成。相比现有顶尖模型,其在时空流畅性、结构稳定性、复杂多主体场景下的指令遵循性以及原生多镜头叙事一致性方面表现卓越。
原文摘要 · Abstract (English)
Notable breakthroughs in diffusion modeling have propelled rapid improvements in video generation, yet current foundational model still face critical challenges in simultaneously balancing prompt following, motion plausibility, and visual quality. In this report, we introduce Seedance 1.0, a high-performance and inference-efficient video foundation generation model that integrates several core technical improvements: (i) multi-source data curation augmented with precision and meaningful video captioning, enabling comprehensive learning across diverse scenarios; (ii) an efficient architecture design with proposed training paradigm, which allows for natively supporting multi-shot generation and jointly learning of both text-to-video and image-to-video tasks. (iii) carefully-optimized post-training approaches leveraging fine-grained supervised fine-tuning, and video-specific RLHF with multi-dimensional reward mechanisms for comprehensive performance improvements; (iv) excellent model acceleration achieving ~10x inference speedup through multi-stage distillation strategies and system-level optimizations. Seedance 1.0 can generate a 5-second video at 1080p resolution only with 41.4 seconds (NVIDIA-L20). Compared to state-of-the-art video generation models, Seedance 1.0 stands out with high-quality and fast video generation having superior spatiotemporal fluidity with structural stability, precise instruction adherence in complex multi-subject contexts, native multi-shot narrative coherence with consistent subject representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。