arXiv:2509.23951cs.CV2025-09被引 134

腾讯发布开源图像生成模型HunyuanImage 3.0,参数达800亿,激活130亿。

HunyuanImage 3.0 Technical Report

  • 基于自回归框架的多模态统一模型,采用思维链设计与渐进式训练策略。
  • 实现800亿总参数、130亿激活参数的MoE结构,生成质量媲美顶尖模型。
  • 代码权重全开源,适合研究者和开发者构建下一代多模态应用。

我们提出HunyuanImage 3.0,一种原生多模态模型,将多模态理解与生成统一于自回归框架中,其图像生成模块已公开。该成果依赖于精细的数据筛选、先进的架构设计、原生思维链机制、渐进式预训练、激进的后训练策略及高效基础设施,支持大规模训练与推理。通过这些改进,我们成功训练了一个总参数超过800亿的Mixture-of-Experts(MoE)模型,推理时每令牌激活约130亿参数,成为迄今最大最强的开源图像生成模型。大量实验表明,自动与人工评估在文本-图像对齐度与视觉质量方面均达到领先水平。通过开放代码与权重,我们希望推动社区利用这一前沿基础模型探索新思路,促进多模态生态发展。所有开源资源可在https://github.com/Tencent-Hunyuan/HunyuanImage-3.0获取。

原文摘要 · Abstract (English)

We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module publicly available. The achievement of HunyuanImage 3.0 relies on several key components, including meticulous data curation, advanced architecture design, a native Chain-of-Thoughts schema, progressive model pre-training, aggressive model post-training, and an efficient infrastructure that enables large-scale training and inference. With these advancements, we successfully trained a Mixture-of-Experts (MoE) model comprising over 80 billion parameters in total, with 13 billion parameters activated per token during inference, making it the largest and most powerful open-source image generative model to date. We conducted extensive experiments and the results of automatic and human evaluation of text-image alignment and visual quality demonstrate that HunyuanImage 3.0 rivals previous state-of-the-art models. By releasing the code and weights of HunyuanImage 3.0, we aim to enable the community to explore new ideas with a state-of-the-art foundation model, fostering a dynamic and vibrant multimodal ecosystem. All open source assets are publicly available at https://github.com/Tencent-Hunyuan/HunyuanImage-3.0

图像生成MoE模型开源多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。