用第一人称视觉驱动虚拟人动作,生成更自然的行走行为。
Moving by Looking: Towards Vision-Driven Avatar Motion Generation
- 通过第一人称视觉输入直接控制虚拟人运动,模拟人类感知方式。
- 训练出的虚拟人能主动避开视野中的障碍物,展现类人运动特征。
- 适合关注虚拟角色自然行为、具身智能与视觉导航的研究者。
我们如何感知世界从根本上影响我们的行动方式,无论是室内移动还是与他人互动。当前的人体动作生成方法忽视了感知与动作之间的关联,使用与人类迥异的任务特定感知机制。本文提出CLOPS,首个仅依赖第一人称视觉感知环境并实现自主导航的虚拟人。由于以视觉为运动主要驱动力,训练面临挑战:现有数据集或缺乏场景上下文,或规模不足。为此,我们解耦低层运动技能与高层视觉控制的学习过程:先在大规模动作捕捉数据上训练运动先验模型;再采用Q-learning训练策略,将第一人称视觉输入映射为运动先验的高层控制指令。实验表明,第一人称视觉可使虚拟人产生类人运动特性,例如主动规避视野中的障碍物。这表明赋予虚拟人类似人类的感官,特别是第一人称视觉,有望训练出行为更像人类的虚拟角色。
原文摘要 · Abstract (English)
The way we perceive the world fundamentally shapes how we move, whether it is how we navigate in a room or how we interact with other humans. Current human motion generation methods, neglect this interdependency and use task-specific ``perception'' that differs radically from that of humans. We argue that the generation of human-like avatar behavior requires human-like perception. Consequently, in this work we present CLOPS, the first human avatar that solely uses egocentric vision to perceive its surroundings and navigate. Using vision as the primary driver of motion however, gives rise to a significant challenge for training avatars: existing datasets have either isolated human motion, without the context of a scene, or lack scale. We overcome this challenge by decoupling the learning of low-level motion skills from learning of high-level control that maps visual input to motion. First, we train a motion prior model on a large motion capture dataset. Then, a policy is trained using Q-learning to map egocentric visual inputs to high-level control commands for the motion prior. Our experiments empirically demonstrate that egocentric vision can give rise to human-like motion characteristics in our avatars. For example, the avatars walk such that they avoid obstacles present in their visual field. These findings suggest that equipping avatars with human-like sensors, particularly egocentric vision, holds promise for training avatars that behave like humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。