让静态重建模型适应第一视角动态视频,消除手部干扰。
Static Scene Reconstruction from Dynamic Egocentric Videos
- 用掩码感知机制抑制注意力层中的动态前景。
- 在HD-EPIC和室内无人机数据集上轨迹误差显著降低。
- 适合需要稳定第一人称3D重建的应用场景。
第一人称视频因相机快速运动和频繁的动态交互,给3D重建带来独特挑战。现有静态重建系统(如MapAnything)在该场景下性能下降,出现严重轨迹漂移和由移动手部导致的“鬼影”几何。本文提出一种鲁棒的流水线,将静态重建主干网络适配于长时第一人称视频。方法引入掩码感知重建机制,在注意力层显式抑制动态前景,防止手部伪影污染静态地图;同时采用分块重建与位姿图拼接策略,保证全局一致性并消除长期漂移。在HD-EPIC和室内无人机数据集上的实验表明,本方法显著降低绝对轨迹误差,并生成视觉清晰的静态几何,有效拓展了基础模型在动态第一人称场景下的能力。
原文摘要 · Abstract (English)
Egocentric videos present unique challenges for 3D reconstruction due to rapid camera motion and frequent dynamic interactions. State-of-the-art static reconstruction systems, such as MapAnything, often degrade in these settings, suffering from catastrophic trajectory drift and "ghost" geometry caused by moving hands. We bridge this gap by proposing a robust pipeline that adapts static reconstruction backbones to long-form egocentric video. Our approach introduces a mask-aware reconstruction mechanism that explicitly suppresses dynamic foreground in the attention layers, preventing hand artifacts from contaminating the static map. Furthermore, we employ a chunked reconstruction strategy with pose-graph stitching to ensure global consistency and eliminate long-term drift. Experiments on HD-EPIC and indoor drone datasets demonstrate that our pipeline significantly improves absolute trajectory error and yields visually clean static geometry compared to naive baselines, effectively extending the capability of foundation models to dynamic first-person scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。