无需特定相机,可统一重建多设备佩戴者全身动作。
OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction
- 用扩散模型实现序列级全身动作重建,避免局部过拟合。
- 在多个真实场景中达到新最优,对遮挡和设备差异鲁棒。
- 适合做穿戴设备行为分析与虚拟人动捕的开发者使用。
消费级第一视角设备的普及为人类行为研究提供了独特视角,但因频繁自遮挡及佩戴者肢体常处于视野外,完整人体三维运动重建仍具挑战。尽管头部与手部轨迹可作为稀疏锚点,现有方法常过度依赖特定硬件光学特性,或依赖昂贵的后期优化,影响动作自然性。本文提出OmniEgoCap,一种统一的扩散框架,可适配多种采集设置。通过从短时窗口估计转向序列级推理,该方法捕捉全局视角并恢复高度、身体比例等不变物理属性,为仅基于头部线索提供关键约束。为实现硬件无关泛化,我们引入几何感知的可见性增强策略,将手部间歇出现视为有原则的几何约束而非缺失数据。模型联合预测时间一致的动作与稳定的身体形态,在公开基准上达到新最优,并在多样真实环境中表现稳健。
原文摘要 · Abstract (English)
The proliferation of commercial egocentric devices offers a unique lens into human behavior, yet reconstructing full-body 3D motion remains difficult due to frequent self-occlusion and the 'out-of-sight' nature of the wearer's limbs. While head and hand trajectories provide sparse anchor points, current methods often overfit to specific hardware optics or rely on expensive, post-hoc optimizations that compromise motion naturalness. In this paper, we present OmniEgoCap, a unified diffusion framework that scales egocentric reconstruction to diverse capture setups. By shifting from short-term windowed estimation to sequence-level inference, our method captures a global perspective and recovers invariant physical attributes, such as height and body proportions, that provide critical constraints for disambiguating head-only cues. To ensure hardware-agnostic generalization, we introduce a geometry-aware visibility augmentation strategy that treats intermittent hand appearances as principled geometric constraints rather than missing data. Our architecture jointly predicts temporally coherent motion and consistent body shape, establishing a new state-of-the-art on public benchmarks and demonstrating robust performance across diverse, in-the-wild environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。