arXiv:2509.21245cs.CVcs.AI2025-09被引 24

支持多模态控制的3D资产生成框架,提升细节精度与生产可用性。

Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets

  • 统一架构融合图像、点云、骨骼等多模态输入进行精细控制。
  • 采用渐进式采样策略,提升对复杂控制信号(如骨骼)的生成能力。
  • 适用于游戏/影视设计,适合需要高可控性的3D内容创作者。

近期3D原生生成模型加速了游戏、影视和设计领域的资产创作。然而,多数方法仍依赖图像或文本条件,缺乏细粒度的跨模态控制,限制了可控性和实际应用。为此,我们提出Hunyuan3D-Omni,一个基于Hunyuan3D 2.1的统一框架,支持点云、体素、边界框和骨骼姿态先验作为条件信号,实现对几何、拓扑和姿态的精确控制。模型采用单一跨模态架构统一处理所有信号,而非为每种模态设置独立分支。通过渐进式、难度感知的采样策略,每个样本仅选择一种控制模态,并偏向更难信号(如骨骼姿态),同时降低易信号(如点云)权重,促进鲁棒的多模态融合并优雅处理缺失输入。实验表明,这些增强控制显著提升了生成精度,支持几何感知变换,并增强了生产工作流的鲁棒性。

原文摘要 · Abstract (English)

Recent advances in 3D-native generative models have accelerated asset creation for games, film, and design. However, most methods still rely primarily on image or text conditioning and lack fine-grained, cross-modal controls, which limits controllability and practical adoption. To address this gap, we present Hunyuan3D-Omni, a unified framework for fine-grained, controllable 3D asset generation built on Hunyuan3D 2.1. In addition to images, Hunyuan3D-Omni accepts point clouds, voxels, bounding boxes, and skeletal pose priors as conditioning signals, enabling precise control over geometry, topology, and pose. Instead of separate heads for each modality, our model unifies all signals in a single cross-modal architecture. We train with a progressive, difficulty-aware sampling strategy that selects one control modality per example and biases sampling toward harder signals (e.g., skeletal pose) while downweighting easier ones (e.g., point clouds), encouraging robust multi-modal fusion and graceful handling of missing inputs. Experiments show that these additional controls improve generation accuracy, enable geometry-aware transformations, and increase robustness for production workflows.

3D生成多模态控制游戏资产

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。