arXiv:2601.19484cs.CV2026-01被引 3

让虚拟人像真人一样感知动态环境并互动,生成更真实的动作。

Dynamic Worlds, Dynamic Humans: Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes

  • 构建视觉-记忆-控制三模块虚拟人架构,模拟人类认知过程。
  • 在动态场景数据集上显著优于现有方法,动作质量与泛化能力提升。
  • 适合做影视动画、游戏角色或元宇宙交互系统的开发者参考。

现实世界中的场景持续动态变化,但现有虚拟人-场景交互生成方法通常将场景视为静态,与实际不符。受世界模型启发,我们提出首个面向动态人-场景交互的认知架构 Dyn-HSI,赋予虚拟人三类类人组件:(1) 视觉(人眼):引入动态场景感知导航,持续感知环境变化并自适应预测下一目标点;(2) 记忆(人脑):设计分层经验记忆,存储并更新训练中积累的经验数据,使模型在推理时可利用先验知识进行上下文感知的动作引导,提升动作质量和泛化性;(3) 控制(人体):采用人-场景交互扩散模型,基于多模态输入生成高保真交互动作。为评估动态场景下的表现,我们扩展现有静态人-场景交互数据集,构建动态基准数据集 Dyn-Scenes。通过大量定性和定量实验验证,Dyn-HSI 在静态与动态场景下均显著优于现有方法,生成高质量的人-场景交互动作。

原文摘要 · Abstract (English)

Scenes are continuously undergoing dynamic changes in the real world. However, existing human-scene interaction generation methods typically treat the scene as static, which deviates from reality. Inspired by world models, we introduce Dyn-HSI, the first cognitive architecture for dynamic human-scene interaction, which endows virtual humans with three humanoid components. (1)Vision (human eyes): we equip the virtual human with a Dynamic Scene-Aware Navigation, which continuously perceives changes in the surrounding environment and adaptively predicts the next waypoint. (2)Memory (human brain): we equip the virtual human with a Hierarchical Experience Memory, which stores and updates experiential data accumulated during training. This allows the model to leverage prior knowledge during inference for context-aware motion priming, thereby enhancing both motion quality and generalization. (3) Control (human body): we equip the virtual human with Human-Scene Interaction Diffusion Model, which generates high-fidelity interaction motions conditioned on multimodal inputs. To evaluate performance in dynamic scenes, we extend the existing static human-scene interaction datasets to construct a dynamic benchmark, Dyn-Scenes. We conduct extensive qualitative and quantitative experiments to validate Dyn-HSI, showing that our method consistently outperforms existing approaches and generates high-quality human-scene interaction motions in both static and dynamic settings.

人机交互动态场景扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。