用稀疏街景实现厘米级3D定位与状态诊断,助力智慧城市建设
SVII-3D: Advancing Roadside Infrastructure Inventory with Decimeter-level 3D Localization and Comprehension from Sparse Street Imagery
- 融合LoRA微调与空间注意力匹配,跨视角关联稀疏图像
- 几何引导优化实现分米级3D定位,误差显著降低
- 引入视觉语言模型自动识别设施运行状态,适合智能运维
自动化构建数字孪生与精确资产清单是智慧城市建设与设施全生命周期管理的关键任务。然而,利用低成本稀疏影像仍面临鲁棒性差、定位不准和细粒度状态理解缺失等挑战。为此,本文提出统一框架SVII-3D,实现资产全链条数字化。首先,将LoRA微调的开放集检测与空间注意力匹配网络融合,稳健关联稀疏视角下的观测;其次,引入几何引导的精修机制,解决结构误差,实现分米级3D定位;第三,突破静态几何映射,通过视觉语言模型代理结合多模态提示,自动诊断细粒度运行状态。实验表明,SVII-3D显著提升识别准确率并最小化定位误差。该框架为高保真基础设施数字化提供了可扩展、低成本的解决方案,有效弥合稀疏感知与自动化智能维护之间的鸿沟。
原文摘要 · Abstract (English)
The automated creation of digital twins and precise asset inventories is a critical task in smart city construction and facility lifecycle management. However, utilizing cost-effective sparse imagery remains challenging due to limited robustness, inaccurate localization, and a lack of fine-grained state understanding. To address these limitations, SVII-3D, a unified framework for holistic asset digitization, is proposed. First, LoRA fine-tuned open-set detection is fused with a spatial-attention matching network to robustly associate observations across sparse views. Second, a geometry-guided refinement mechanism is introduced to resolve structural errors, achieving precise decimeter-level 3D localization. Third, transcending static geometric mapping, a Vision-Language Model agent leveraging multi-modal prompting is incorporated to automatically diagnose fine-grained operational states. Experiments demonstrate that SVII-3D significantly improves identification accuracy and minimizes localization errors. Consequently, this framework offers a scalable, cost-effective solution for high-fidelity infrastructure digitization, effectively bridging the gap between sparse perception and automated intelligent maintenance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。