arXiv:2510.11650cs.CV2025-10SIGGRAPH被引 7

用现有模型自动生成无限多样且可控的3D人像数据,低成本实现高质量虚拟角色生成。

InfiniHuman: Infinite 3D Human Creation with Precise Control

  • 利用视觉语言模型自动构建多模态3D人像数据集
  • 生成11.1万种身份,支持文本、姿态与服装精细控制
  • 适合需要海量多样化角色的元宇宙与影视制作场景

生成真实且可控制的3D人类角色是长期挑战,尤其在种族、年龄、服饰与体型等属性上需广泛覆盖。传统大规模人体数据采集与标注成本高昂,且规模与多样性受限。本文核心问题为:能否通过蒸馏现有基础模型生成理论上无限、丰富标注的3D人体数据?我们提出InfiniHuman框架,协同蒸馏多种模型,以极低成本实现理论上无限制的可扩展数据生成。我们设计InfiniHumanData,一个全自动多模态数据生成管道,融合视觉-语言与图像生成模型,构建大规模数据集。用户研究显示,自动生成的身份与扫描渲染结果难以区分。InfiniHumanData包含111,000个身份,涵盖前所未有的多样性。每个身份配有细粒度文本描述、多视角RGB图像、详细服装图像及SMPL体形参数。基于此数据集,我们提出InfiniHumanGen,一种基于扩散模型的生成流程,支持文本、体形与服装资产条件输入。该方法实现快速、逼真且精准可控的虚拟角色生成。大量实验表明,在视觉质量、生成速度与可控性方面均显著优于现有方法。本方案通过实用且经济的方式,实现高质量、细粒度控制的无限规模角色生成。相关代码、数据集与模型将公开发布于https://yuxuan-xue.com/infini-human。

原文摘要 · Abstract (English)

Generating realistic and controllable 3D human avatars is a long-standing challenge, particularly when covering broad attribute ranges such as ethnicity, age, clothing styles, and detailed body shapes. Capturing and annotating large-scale human datasets for training generative models is prohibitively expensive and limited in scale and diversity. The central question we address in this paper is: Can existing foundation models be distilled to generate theoretically unbounded, richly annotated 3D human data? We introduce InfiniHuman, a framework that synergistically distills these models to produce richly annotated human data at minimal cost and with theoretically unlimited scalability. We propose InfiniHumanData, a fully automatic pipeline that leverages vision-language and image generation models to create a large-scale multi-modal dataset. User study shows our automatically generated identities are undistinguishable from scan renderings. InfiniHumanData contains 111K identities spanning unprecedented diversity. Each identity is annotated with multi-granularity text descriptions, multi-view RGB images, detailed clothing images, and SMPL body-shape parameters. Building on this dataset, we propose InfiniHumanGen, a diffusion-based generative pipeline conditioned on text, body shape, and clothing assets. InfiniHumanGen enables fast, realistic, and precisely controllable avatar generation. Extensive experiments demonstrate significant improvements over state-of-the-art methods in visual quality, generation speed, and controllability. Our approach enables high-quality avatar generation with fine-grained control at effectively unbounded scale through a practical and affordable solution. We will publicly release the automatic data generation pipeline, the comprehensive InfiniHumanData dataset, and the InfiniHumanGen models at https://yuxuan-xue.com/infini-human.

3D生成虚拟角色可控生成数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。