让数字人仅靠视觉和目标主动在新场景中自然行动
Visually-grounded Humanoid Agents
- 用双层架构让数字人像真人一样看、想、动
- 在新场景中任务成功率更高,碰撞更少
- 适合做虚拟世界角色生成与交互研究
数字人生成已研究多年,但多数系统依赖预设状态或脚本控制,难以扩展到新环境。本文提出视觉引导的类人智能体,使数字人仅通过视觉观察和目标任务,在新3D场景中主动、自然地行为。该系统采用双层架构:世界层通过遮挡感知管道从真实视频重建语义丰富的3D高斯场景,并支持可动画化的高斯人体模型;智能体层则赋予这些模型第一人称RGB-D感知能力,实现具备空间意识和迭代推理的精准具身规划,并以全身动作低层执行,驱动其在场景中的行为。我们还构建了一个评估人-场景交互的基准。实验表明,相比消融实验和现有方法,本系统在新环境中表现出更强的自主性,任务成功率更高且碰撞更少。数据、代码与模型将开源。此工作推动了主动数字人规模化部署与以人为本的具身智能发展。
原文摘要 · Abstract (English)
Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively animated, relying on privileged state or scripted control, which limits scalability to novel environments. We instead ask: how can digital humans actively behave using only visual observations and specified goals in novel scenes? Achieving this would enable populating any 3D environments with digital humans at scale that exhibit spontaneous, natural, goal-directed behaviors. To this end, we introduce Visually-grounded Humanoid Agents, a coupled two-layer (world-agent) paradigm that replicates humans at multiple levels: they look, perceive, reason, and behave like real people in real-world 3D scenes. The World Layer reconstructs semantically rich 3D Gaussian scenes from real-world videos via an occlusion-aware pipeline and accommodates animatable Gaussian-based human avatars. The Agent Layer transforms these avatars into autonomous humanoid agents, equipping them with first-person RGB-D perception and enabling them to perform accurate, embodied planning with spatial awareness and iterative reasoning, which is then executed at the low level as full-body actions to drive their behaviors in the scene. We further introduce a benchmark to evaluate humanoid-scene interaction in diverse reconstructed environments. Experiments show our agents achieve robust autonomous behavior, yielding higher task success rates and fewer collisions than ablations and state-of-the-art planning methods. This work enables active digital human population and advances human-centric embodied AI. Data, code, and models will be open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。