arXiv:2512.22808cs.CVcs.AI2025-12被引 8

基于第一人称视频生成3D人体反应动作,实现时空对齐与实时生成。

EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation

  • 用向量量化VAE压缩动作,结合GPT生成反应序列。
  • 在真实数据集上生成动作的时空一致性提升47%以上。
  • 适合虚拟现实、智能助手等需要自然交互的应用场景。

人类对第一人称视觉输入表现出适应性、上下文敏感的反应行为。然而,由于严格因果生成和精确三维空间对齐的双重需求,从第一人称视频中建模此类反应仍具挑战性。为此,我们构建了人类反应数据集(HRD),通过建立空间对齐的第一人称视频-反应数据集,解决现有数据集(如ViMo)中存在的显著空间不一致问题,例如动态动作常与固定视角视频配对。基于HRD,我们提出EgoReAct,首个能从第一人称视频流中实时生成三维对齐人体反应动作的自回归框架。首先通过向量量化变分自编码器将反应动作压缩至紧凑且表达性强的潜在空间,再训练生成式预训练变换器根据视觉输入生成反应。EgoReAct引入三维动态特征,包括度量深度和头部运动,在生成过程中有效增强空间定位能力。大量实验表明,相比以往方法,EgoReAct在真实性、空间一致性及生成效率方面均有显著提升,同时保持生成过程的严格因果性。代码、模型与数据将在录用后发布。

原文摘要 · Abstract (English)

Humans exhibit adaptive, context-sensitive responses to egocentric visual input. However, faithfully modeling such reactions from egocentric video remains challenging due to the dual requirements of strictly causal generation and precise 3D spatial alignment. To tackle this problem, we first construct the Human Reaction Dataset (HRD) to address data scarcity and misalignment by building a spatially aligned egocentric video-reaction dataset, as existing datasets (e.g., ViMo) suffer from significant spatial inconsistency between the egocentric video and reaction motion, e.g., dynamically moving motions are always paired with fixed-camera videos. Leveraging HRD, we present EgoReAct, the first autoregressive framework that generates 3D-aligned human reaction motions from egocentric video streams in real-time. We first compress the reaction motion into a compact yet expressive latent space via a Vector Quantised-Variational AutoEncoder and then train a Generative Pre-trained Transformer for reaction generation from the visual input. EgoReAct incorporates 3D dynamic features, i.e., metric depth, and head dynamics during the generation, which effectively enhance spatial grounding. Extensive experiments demonstrate that EgoReAct achieves remarkably higher realism, spatial consistency, and generation efficiency compared with prior methods, while maintaining strict causality during generation. We will release code, models, and data upon acceptance.

第一人称视频3D动作生成实时生成空间对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。