arXiv:2506.05546cs.CV2025-06CVPR

将2D运动分割结果融合进分层辐射场,实现动态场景的3D运动分割

Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos

  • 用分层辐射场融合2D运动分割预测,构建3D动态结构
  • 在EPIC-KITCHENS数据集上,3D分割精度比2D基线提升18.6% mAP
  • 适用于长时序、高动态的第一视角视频分析,适合视觉理解研究者

计算机视觉仍以2D技术为主,3D视觉应用受限。尽管神经辐射场等进展使3D方法能通过融合多视角2D输出并去噪来提升性能,但现有方法在动态场景中效果不佳。本文针对第一人称视频中的动态物体分割问题,提出分层运动融合(Layered Motion Fusion):将基于2D模型的运动分割结果注入分层辐射场。为应对长时序动态视频带来的几何结构捕捉难题,引入测试时精炼机制,聚焦关键帧以降低复杂度。该策略实现了运动融合与精炼的协同,使3D模型的分割精度显著超越2D基线,在EPIC-KITCHENS数据集上达到18.6%的mAP提升,证明了3D方法在真实动态场景中增强2D分析的可行性。

原文摘要 · Abstract (English)

Computer vision is largely based on 2D techniques, with 3D vision still relegated to a relatively narrow subset of applications. However, by building on recent advances in 3D models such as neural radiance fields, some authors have shown that 3D techniques can at last improve outputs extracted from independent 2D views, by fusing them into 3D and denoising them. This is particularly helpful in egocentric videos, where the camera motion is significant, but only under the assumption that the scene itself is static. In fact, as shown in the recent analysis conducted by EPIC Fields, 3D techniques are ineffective when it comes to studying dynamic phenomena, and, in particular, when segmenting moving objects. In this paper, we look into this issue in more detail. First, we propose to improve dynamic segmentation in 3D by fusing motion segmentation predictions from a 2D-based model into layered radiance fields (Layered Motion Fusion). However, the high complexity of long, dynamic videos makes it challenging to capture the underlying geometric structure, and, as a result, hinders the fusion of motion cues into the (incomplete) scene geometry. We address this issue through test-time refinement, which helps the model to focus on specific frames, thereby reducing the data complexity. This results in a synergy between motion fusion and the refinement, and in turn leads to segmentation predictions of the 3D model that surpass the 2D baseline by a large margin. This demonstrates that 3D techniques can enhance 2D analysis even for dynamic phenomena in a challenging and realistic setting.

3D分割动态场景第一人称视频辐射场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。