arXiv:2511.18870cs.CV2025-11被引 107

83亿参数的开源视频生成模型,能在消费级显卡上高效运行。

HunyuanVideo 1.5 Technical Report

  • 采用选择性滑动块注意力机制,提升生成效率与连贯性。
  • 在多时长多分辨率下实现顶尖视觉质量,超越现有开源模型。
  • 适合想低成本尝试高质量视频生成的研究者和创作者。

我们提出HunyuanVideo 1.5,一个仅含83亿参数的轻量级开源视频生成模型,在消费级GPU上实现高效推理的同时达到当前最优的视觉质量和运动连贯性。该成果基于精心的数据清洗、先进的DiT架构(含选择性与滑动块注意力SSTA)、通过字形感知文本编码增强的双语理解能力、渐进式预训练与后训练流程,以及高效的视频超分网络。构建统一框架,支持跨时长、跨分辨率的文本到视频与图像到视频生成。大量实验证明,该紧凑模型在开源领域树立新标杆。代码与模型权重已公开发布于https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5,为视频创作与研究提供高性能基础,降低技术门槛。

原文摘要 · Abstract (English)

We present HunyuanVideo 1.5, a lightweight yet powerful open-source video generation model that achieves state-of-the-art visual quality and motion coherence with only 8.3 billion parameters, enabling efficient inference on consumer-grade GPUs. This achievement is built upon several key components, including meticulous data curation, an advanced DiT architecture featuring selective and sliding tile attention (SSTA), enhanced bilingual understanding through glyph-aware text encoding, progressive pre-training and post-training, and an efficient video super-resolution network. Leveraging these designs, we developed a unified framework capable of high-quality text-to-video and image-to-video generation across multiple durations and resolutions. Extensive experiments demonstrate that this compact and proficient model establishes a new state-of-the-art among open-source video generation models. By releasing the code and model weights, we provide the community with a high-performance foundation that lowers the barrier to video creation and research, making advanced video generation accessible to a broader audience. All open-source assets are publicly available at https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5.

视频生成扩散模型开源模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。