arXiv:2503.08165cs.CV2025-03被引 1

用文字或图片生成可动画化的3D人像,支持精细定制。

Multimodal Generation of Animatable 3D Human Models with AvatarForge

  • 结合大模型与现成3D生成器,实现精准控制人体细节。
  • 在文本和图像生成任务中均超越现有方法,效果更优。
  • 适合艺术创作与动画设计,支持持续优化与迭代。

我们提出AvatarForge,一种基于AI驱动的程序化生成框架,可从文本或图像输入生成可动画化的3D人类角色。尽管基于扩散的方法在通用3D物体生成上取得进展,但面对人类体型、姿态的高度复杂性与多样性,以及高质量数据稀缺的问题,仍难以生成高精度且可定制的人类角色。此外,现有方法在角色动画方面也面临挑战。AvatarForge通过结合大语言模型的常识推理与现成的3D人体生成器,实现了对身体与面部细节的细粒度控制。相比依赖预训练数据集且缺乏个体特征精确控制的扩散模型,AvatarForge提供更灵活的交互式建模流程,并配备自动验证系统,支持生成结果的持续优化,显著提升准确性与个性化程度。评估显示,该方法在文本到角色与图像到角色生成任务中均优于当前最优方法,展现出强大的艺术创作与动画应用潜力。

原文摘要 · Abstract (English)

We introduce AvatarForge, a framework for generating animatable 3D human avatars from text or image inputs using AI-driven procedural generation. While diffusion-based methods have made strides in general 3D object generation, they struggle with high-quality, customizable human avatars due to the complexity and diversity of human body shapes, poses, exacerbated by the scarcity of high-quality data. Additionally, animating these avatars remains a significant challenge for existing methods. AvatarForge overcomes these limitations by combining LLM-based commonsense reasoning with off-the-shelf 3D human generators, enabling fine-grained control over body and facial details. Unlike diffusion models which often rely on pre-trained datasets lacking precise control over individual human features, AvatarForge offers a more flexible approach, bringing humans into the iterative design and modeling loop, with its auto-verification system allowing for continuous refinement of the generated avatars, and thus promoting high accuracy and customization. Our evaluations show that AvatarForge outperforms state-of-the-art methods in both text- and image-to-avatar generation, making it a versatile tool for artistic creation and animation.

3D生成多模态可动画角色设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。