arXiv:2506.04606cs.CV2025-06被引 1

用AI助手生成可动3D人像,支持文本和照片输入。

SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents

  • 通过视觉语言模型驱动的自动验证循环,实现精细控制。
  • 生成的3D人像身份一致、身体结构合理、支持动作变形。
  • 适合需要定制化可动画角色的创作者或游戏开发者。

SmartAvatar是一种基于视觉-语言-智能体的框架,仅需一张照片或文本提示即可生成完全绑定、可动画的3D人像。尽管扩散模型在通用3D物体生成上取得进展,但在人物身份、体型控制和动画就绪性方面仍存在挑战。SmartAvatar结合大视觉语言模型(VLM)的常识推理能力与现成参数化人体生成器,实现高质量、可定制的人像生成。其核心创新在于自主验证循环:智能体渲染草图,评估面部相似性、解剖合理性与提示匹配度,并迭代调整生成参数直至收敛。该交互式AI引导优化过程使用户可通过自然语言对话逐步完善人像。相比依赖静态预训练数据集且灵活性有限的扩散模型,SmartAvatar将用户纳入建模流程,通过LLM驱动的程序化生成与验证系统实现持续改进。生成的人像完全绑定,支持姿态操控,保持身份与外观一致性,适用于下游动画与交互应用。定量基准与用户研究显示,SmartAvatar在网格重建质量、身份保真度、属性准确性和动画就绪性方面均优于近期文本与图像驱动的人像生成系统,是可在消费级硬件上运行的高实用性工具。

原文摘要 · Abstract (English)

SmartAvatar is a vision-language-agent-driven framework for generating fully rigged, animation-ready 3D human avatars from a single photo or textual prompt. While diffusion-based methods have made progress in general 3D object generation, they continue to struggle with precise control over human identity, body shape, and animation readiness. In contrast, SmartAvatar leverages the commonsense reasoning capabilities of large vision-language models (VLMs) in combination with off-the-shelf parametric human generators to deliver high-quality, customizable avatars. A key innovation is an autonomous verification loop, where the agent renders draft avatars, evaluates facial similarity, anatomical plausibility, and prompt alignment, and iteratively adjusts generation parameters for convergence. This interactive, AI-guided refinement process promotes fine-grained control over both facial and body features, enabling users to iteratively refine their avatars via natural-language conversations. Unlike diffusion models that rely on static pre-trained datasets and offer limited flexibility, SmartAvatar brings users into the modeling loop and ensures continuous improvement through an LLM-driven procedural generation and verification system. The generated avatars are fully rigged and support pose manipulation with consistent identity and appearance, making them suitable for downstream animation and interactive applications. Quantitative benchmarks and user studies demonstrate that SmartAvatar outperforms recent text- and image-driven avatar generation systems in terms of reconstructed mesh quality, identity fidelity, attribute accuracy, and animation readiness, making it a versatile tool for realistic, customizable avatar creation on consumer-grade hardware.

3D生成人像建模AI代理动画适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。