arXiv:2607.18153cs.CV2026-07中稿 · ICRA

融合多模态信息实现精准动态物体分割,提升边界一致性。

Robust Multimodal Dynamic Object Segmentation

论文配图:Robust Multimodal Dynamic Object Segmentation
图 1 · 摘自论文原文
  • 结合2D轨迹、3D重建与语义信息进行联合分类
  • 在多个数据集上达到领先性能,显著改善分割边界质量
  • 适合需要高精度动态物体识别的自动驾驶场景

动态物体分割在静态场景重建等视觉任务中至关重要。现有基于光流的方法难以保证物体边界处静态/动态分割的一致性,而基于3D重建的方法对重建误差极为敏感。为此,我们提出一种融合多模态线索(包括2D点轨迹、3D重建和语义信息)的动态物体分割框架,可生成精确且完整的动态掩码。设计了结合Transformer与特征聚类聚合模块的网络,对多模态特征轨迹进行静态/动态分类,能自适应选择主导特征类型,并缓解特征退化影响。此外,引入基于点查询的SAM后处理方法,可有效处理单个掩码中的多个物体。大量实验表明,该方法在动态物体分割与静态场景重建任务中均达当前最优性能。

原文摘要 · Abstract (English)

Dynamic object segmentation plays a critical role in many visual applications such as static scene reconstruction from dynamic videos. However, existing optical flow-based methods fail to ensure consistent static/dynamic segmentation along object boundaries, while 3D reconstruction-based approaches are highly sensitive to reconstruction errors. To address these limitations, we present a dynamic object segmentation framework that can generate both precise and complete dynamic masks by integrating multimodal cues including 2D point tracks, 3D reconstruction, and semantic information. We design a network combining Transformer architectures with feature clustering aggregation modules to perform static/dynamic classification of multimodal feature trajectories. It enables the model to adaptively determine which type of feature should dominate based on the characteristics of each scene, while also mitigating the impact of feature degradation. Additionally, we introduce a novel point-query-based SAM post-processing method capable of handling multiple objects within a single mask. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in both dynamic object segmentation and static scene reconstruction tasks.

动态分割多模态3D重建语义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。