用3D几何信息提升街景图像中物体的跨视角识别与追踪能力。
Joint 2D-3D Segmentation and Association in Street-level Imaging

- 融合2D视觉与3D几何推理,实现跨视角物体关联。
- 在复杂城市场景下,身份保持性能提升22%。
- 适用于大规模街景建模,支持多种物体类型。
准确理解街景图像对大规模城市测绘和空间数字孪生(SDT)环境构建至关重要。本文提出一种统一框架,实现2D-3D联合分割与关联,将视觉语义与多视角几何推理相结合。不同于依赖连续帧进行时序追踪的传统方法,本方法结合零样本检测与分割,以及基于运动恢复结构(SfM)的重建,建立稳定的跨视角对应关系。采用3D驱动的关联机制替代传统2D多目标追踪,利用几何一致性在远基线视角和不同成像条件下保持物体身份。通过融合2D纹理线索与全局3D上下文,该流程适用于可扩展的街景处理,可适配多种物体类型。实验表明,在真实数据序列中覆盖度显著提升,身份保留更鲁棒,相较于当前最先进的2D仅追踪方法,在挑战性城市场景中性能提升22%。
原文摘要 · Abstract (English)
Accurate interpretation of street-level imagery is essential for large-scale urban mapping and the creation of Spatial Digital Twin (SDT) environments. This work presents a unified framework for joint 2D-3D segmentation and association that integrates visual semantics with multi-view geometric reasoning. Unlike conventional approaches that rely heavily on sequential frames for temporal tracking, our method leverages zero-shot detection and segmentation together with structure-from-motion reconstruction to establish stable cross-view correspondences. A 3D-driven association mechanism replaces traditional 2D multi-object tracking, using geometric consistency to guide identity preservation across wide-baseline viewpoints and varying imaging conditions. By combining 2D texture cues with global 3D context, the proposed pipeline is well-suited for scalable street-level processing and can be used for a variety of object types. Experiments demonstrate substantially improved coverage of ground-truth sequences and more robust identity retention compared to state-of-the-art 2D-only tracking methods, achieving a 22% performance gain in challenging urban scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。