arXiv:2604.19105cs.CV2026-04被引 1

让第一人称视角动作生成更自然,靠分步推理与扩散模型协同实现

EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation

论文配图:EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
图 1 · 摘自论文原文
  • 分两阶段:先用视觉语言模型提炼动作指令,再用扩散模型生成流畅动作
  • 在多个数据集上超越现有方法,动作更符合语义且物理合理
  • 适合做具身智能、人机交互和虚拟角色动画的研究者参考

在动态环境中忠实建模人类行为是具身智能的基础挑战。尽管条件化动作合成已取得显著进展,第一人称视角(egocentric)的动作生成仍因第一人称感知的固有复杂性而研究不足。本文研究第一人称视觉-语言(Ego-VL)动作生成任务,即根据第一人称视觉输入和自然语言指令联合生成3D人体动作。我们识别出关键的“推理-生成耦合”挑战:同时优化语义推理与运动建模导致梯度冲突,系统性降低多模态对齐精度和动作质量。为此,我们提出分层生成框架EgoMotion。受生物认知与运动控制解耦启发,EgoMotion分为两个阶段:在认知推理阶段,视觉语言模型(VLM)将多模态输入映射到离散动作基元的结构化空间,强制VLM获得目标一致的表征,有效弥合高层感知理解与底层动作执行间的语义鸿沟;在运动生成阶段,这些学习到的表征作为表达性强的条件信号,驱动基于扩散的运动生成器。通过在连续潜在空间中迭代去噪,生成器合成物理合理且时间连贯的动作轨迹。大量评估表明,EgoMotion达到当前最优性能,生成的动作序列在语义对齐性和运动质量上均优于现有方法。

原文摘要 · Abstract (English)

Faithfully modeling human behavior in dynamic environments is a foundational challenge for embodied intelligence. While conditional motion synthesis has achieved significant advances, egocentric motion generation remains largely underexplored due to the inherent complexity of first-person perception. In this work, we investigate Egocentric Vision-Language (Ego-VL) motion generation. This task requires synthesizing 3D human motion conditioned jointly on first-person visual observations and natural language instructions. We identify a critical \textit{reasoning-generation entanglement} challenge: the simultaneous optimization of semantic reasoning and kinematic modeling introduces gradient conflicts. These conflicts systematically degrade the fidelity of multimodal grounding and motion quality. To address this challenge, we propose a hierarchical generative framework \textbf{EgoMotion}. Inspired by the biological decoupling of cognitive reasoning and motor control, EgoMotion operates in two stages. In the Cognitive Reasoning stage, A vision-language model (VLM) projects multimodal inputs into a structured space of discrete motion primitives. This forces the VLM to acquire goal-consistent representations, effectively bridging the semantic gap between high-level perceptual understanding and low-level action execution. In the Motion Generation stage, these learned representations serve as expressive conditioning signals for a diffusion-based motion generator. By performing iterative denoising within a continuous latent space, the generator synthesizes physically plausible and temporally coherent trajectories. Extensive evaluations demonstrate that EgoMotion achieves state-of-the-art performance, and produces motion sequences that are both semantically grounded and kinematically superior to existing approaches.

动作生成视觉语言扩散模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。