用文本生成高精度3D人物与物体互动,解决提示遵循不准的问题。
Hoi3DGen: Generating High-Quality Human-Object-Interactions in 3D
- 基于多模态大模型构建高质量交互数据,训练文本到3D的生成管道
- 在文本一致性上提升4-15倍,3D模型质量提升3-7倍
- 支持多种类别和交互类型,适合游戏、AR/XR应用开发
从文本建模与生成3D人-物交互对AR、XR和游戏应用至关重要。现有方法常依赖文生图模型的分数蒸馏,但受限于高质量交互数据稀缺,导致结果存在‘双面人’问题且不忠实于文本提示。我们提出Hoi3DGen框架,可生成精确遵循输入描述的高质量带纹理3D网格。首先利用多模态大语言模型构建真实、高质量的交互数据,进而建立完整的文本到3D生成流水线,在交互保真度上实现数量级提升。相比基线方法,该方法在文本一致性上提升4-15倍,3D模型质量提升3-7倍,展现出对多样类别和交互类型的强泛化能力,同时保持高水准3D生成效果。
原文摘要 · Abstract (English)
Modeling and generating 3D human-object interactions from text is crucial for applications in AR, XR, and gaming. Existing approaches often rely on score distillation from text-to-image models, but their results suffer from the Janus problem and do not follow text prompts faithfully due to the scarcity of high-quality interaction data. We introduce Hoi3DGen, a framework that generates high-quality textured meshes of human-object interaction that follow the input interaction descriptions precisely. We first curate realistic and high-quality interaction data leveraging multimodal large language models, and then create a full text-to-3D pipeline, which achieves orders-of-magnitude improvements in interaction fidelity. Our method surpasses baselines by 4-15x in text consistency and 3-7x in 3D model quality, exhibiting strong generalization to diverse categories and interaction types, while maintaining high-quality 3D generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。