arXiv:2602.18432cs.CV2026-02被引 3

让虚拟人物实时感知用户位置并自然互动

SARAH: Spatially Aware Real-time Agentic Humans

  • 用因果Transformer+流匹配模型实现空间感知的实时动作生成
  • 在Embody 3D数据集上达300+帧率,速度是非因果方法3倍
  • 支持用户动态调节眼神交流强度,适合虚拟人和VR应用

随着具身智能体在虚拟现实、远程呈现和数字人类应用中的重要性提升,其运动需超越语音对齐的手势:智能体应转向用户、响应用户移动,并保持自然注视。现有方法缺乏这种空间感知能力。本文提出首个实时、完全因果的空间感知对话动作生成方法,可部署于流式VR头显。给定用户位置与双人音频输入,该方法生成与语音对齐且面向用户的全身动作。架构结合因果Transformer-based VAE与交错潜码以支持流式推理,以及基于用户轨迹和音频条件的流匹配模型。为支持不同注视偏好,引入带有无分类器引导的眼球注视评分机制:模型从数据中学习自然空间对齐,用户可在推理时调整眼神接触强度。在Embody 3D数据集上,本方法实现超过300 FPS的最优动作质量,较非因果基线快3倍,精准捕捉自然对话中的细微空间动态。我们已在真实VR系统中验证,实现空间感知对话智能体的实时部署。

原文摘要 · Abstract (English)

As embodied agents become central to VR, telepresence, and digital human applications, their motion must go beyond speech-aligned gestures: agents should turn toward users, respond to their movement, and maintain natural gaze. Current methods lack this spatial awareness. We close this gap with the first real-time, fully causal method for spatially-aware conversational motion, deployable on a streaming VR headset. Given a user's position and dyadic audio, our approach produces full-body motion that aligns gestures with speech while orienting the agent according to the user. Our architecture combines a causal transformer-based VAE with interleaved latent tokens for streaming inference and a flow matching model conditioned on user trajectory and audio. To support varying gaze preferences, we introduce a gaze scoring mechanism with classifier-free guidance to decouple learning from control: the model captures natural spatial alignment from data, while users can adjust eye contact intensity at inference time. On the Embody 3D dataset, our method achieves state-of-the-art motion quality at over 300 FPS -- 3x faster than non-causal baselines -- while capturing the subtle spatial dynamics of natural conversation. We validate our approach on a live VR system, bringing spatially-aware conversational agents to real-time deployment. Please see https://evonneng.github.io/sarah/ for details.

虚拟人实时生成空间感知VR交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。