arXiv:2502.08556cs.CVcs.AI2025-02IJCAI被引 15

提出人类中心基础模型四类框架,统一感知、生成与智能行为建模。

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling

  • 按感知、生成、统一、智能代理四类构建人类中心模型体系
  • 整合多模态2D/3D理解与高保真内容生成能力
  • 适合数字人、具身智能研究者参考,推动智能体发展

人类理解与生成是建模数字人和人形实体的关键。近期,受通用模型(如大语言模型、视觉模型)成功启发,人类中心基础模型(HcFMs)应运而生,将多样化的以人为中心的任务统一于单一框架中,超越传统任务专用方法。本文通过提出分类体系,系统梳理当前研究,将其分为四类:(1) 人类中心感知基础模型,用于捕捉多模态2D/3D的细粒度特征;(2) 人类中心AIGC基础模型,生成高质量、多样化的相关人类内容;(3) 统一感知与生成模型,融合理解与合成能力;(4) 人类中心智能体基础模型,延伸至学习类人智能与交互行为,支持人形实体任务。综述了前沿技术,讨论新兴挑战与未来方向,旨在为数字人及具身智能建模研究提供路线图。

原文摘要 · Abstract (English)

Human understanding and generation are critical for modeling digital humans and humanoid embodiments. Recently, Human-centric Foundation Models (HcFMs) inspired by the success of generalist models, such as large language and vision models, have emerged to unify diverse human-centric tasks into a single framework, surpassing traditional task-specific approaches. In this survey, we present a comprehensive overview of HcFMs by proposing a taxonomy that categorizes current approaches into four groups: (1) Human-centric Perception Foundation Models that capture fine-grained features for multi-modal 2D and 3D understanding. (2) Human-centric AIGC Foundation Models that generate high-fidelity, diverse human-related content. (3) Unified Perception and Generation Models that integrate these capabilities to enhance both human understanding and synthesis. (4) Human-centric Agentic Foundation Models that extend beyond perception and generation to learn human-like intelligence and interactive behaviors for humanoid embodied tasks. We review state-of-the-art techniques, discuss emerging challenges and future research directions. This survey aims to serve as a roadmap for researchers and practitioners working towards more robust, versatile, and intelligent digital human and embodiments modeling.

数字人基础模型具身智能AIGC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。