arXiv:2608.02711cs.CV2026-08

统一3D生成理解与编辑,支持文本到3D和指令式修改。

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

论文配图:Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
图 1 · 摘自论文原文
  • 用统一架构整合3D理解、生成与编辑,共享模型参数。
  • 构建8700万规模3D多模态数据集,含1200万编辑样本。
  • 支持结构保持的局部编辑,适合工业设计与内容创作。

近期图像生成进展表明,统一多模态模型在理解、生成与编辑方面具有潜力。然而,统一3D建模受限于稀缺的多模态数据,尤其缺乏大规模且几何一致的编辑数据。为此,我们提出Hunyuan3D-Buffalo 1.0,一个统一框架,支持3D理解、文本到3D生成、指令引导的3D编辑及文本驱动的部件生成。为实现可扩展训练,我们构建了8700万规模的3D多模态语料库,包括2500万理解样本、5000万文本-3D配对和1200万编辑样本(基于Nano3D-v2生成)。架构上,结合Hunyuan3D-VLM用于语义、结构与空间理解,以及Hunyuan3D DiT实现高保真3D合成。VLM提供生成的多模态语义条件,而编辑与部件生成还额外以源对象表示作为扩散过程的条件,以保持其整体结构与未编辑区域。大量实验表明,Hunyuan3D-Buffalo 1.0在文本到3D生成与3D编辑基准上达到最先进或领先性能,同时展现出强大的理解与部件生成能力。分析进一步显示,生成与理解能力共同提升编辑效果,验证了统一3D多模态训练的有效性。

原文摘要 · Abstract (English)

Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/

3D生成多模态编辑扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。