Hunyuan3D 1.0实现文本与图像到3D的快速生成,兼顾速度与质量。
Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
- 两阶段架构:先多视角扩散生成图像,再前馈重建3D模型。
- 标准版生成仅需约11秒(4秒+7秒),支持高质量3D资产输出。
- 统一框架兼容文本与图像输入,适合需要高效3D创作的设计师。
尽管3D生成模型显著提升了艺术家的工作流程,现有基于扩散的3D生成模型仍存在生成速度慢、泛化能力差的问题。为此,我们提出一种名为Hunyuan3D 1.0的两阶段方法,包含轻量版和标准版,均支持文本与图像条件生成。第一阶段采用多视图扩散模型,约4秒内生成多视角RGB图像,捕捉3D资产的丰富细节,将任务从单视图重构转为多视图重构。第二阶段引入前馈重建模型,约7秒内根据生成的多视图图像精准重建3D资产,学习处理多视图扩散带来的噪声与不一致性,并利用条件图像信息高效恢复3D结构。该框架集成文本到图像模型Hunyuan-DiT,实现文本与图像条件3D生成的统一。标准版参数量比轻量版多3倍,且超过其他现有模型。Hunyuan3D 1.0在速度与质量间取得出色平衡,显著缩短生成时间同时保持产出资产的质量与多样性。
原文摘要 · Abstract (English)
While 3D generative models have greatly improved artists' workflows, the existing diffusion models for 3D generation suffer from slow generation and poor generalization. To address this issue, we propose a two-stage approach named Hunyuan3D 1.0 including a lite version and a standard version, that both support text- and image-conditioned generation. In the first stage, we employ a multi-view diffusion model that efficiently generates multi-view RGB in approximately 4 seconds. These multi-view images capture rich details of the 3D asset from different viewpoints, relaxing the tasks from single-view to multi-view reconstruction. In the second stage, we introduce a feed-forward reconstruction model that rapidly and faithfully reconstructs the 3D asset given the generated multi-view images in approximately 7 seconds. The reconstruction network learns to handle noises and in-consistency introduced by the multi-view diffusion and leverages the available information from the condition image to efficiently recover the 3D structure. Our framework involves the text-to-image model, i.e., Hunyuan-DiT, making it a unified framework to support both text- and image-conditioned 3D generation. Our standard version has 3x more parameters than our lite and other existing model. Our Hunyuan3D 1.0 achieves an impressive balance between speed and quality, significantly reducing generation time while maintaining the quality and diversity of the produced assets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。