用语言指令控制无人机3D重建,实现智能巡检
From Flight to Insight: Semantic 3D Reconstruction for Aerial Inspection via Gaussian Splatting and Language-Guided Segmentation

- 结合3D高斯泼溅与语言引导分割,生成语义可理解的三维模型
- 通过语言提示生成热图,再用SAM2精修2D分割,准确率提升显著
- 适合需要快速语义分析的无人机巡检、基建监测场景
高保真3D重建对基础设施监测、结构评估和环境调查等航拍任务至关重要。传统摄影测量虽能建模几何结构,但缺乏语义解释能力,限制了自动化巡检的应用。近年来神经渲染与3D高斯泼溅(3DGS)技术可实现高效逼真的重建,但仍缺乏场景级理解。本文提出一种基于无人机的管线,扩展Feature-3DGS以支持语言引导的3D分割。利用基于LSeg的特征场与CLIP嵌入生成响应语言提示的热图,经阈值处理得到粗略分割,并以得分最高的点作为提示,对新视角渲染图使用SAM或SAM2进行精细2D分割。实验表明,不同特征场主干(CLIP-LSeg、SAM、SAM2)在大型户外环境中捕捉有意义结构的能力各异。该混合方法实现了对逼真3D重建的灵活语言交互,为语义化航拍巡检与场景理解开辟新可能。
原文摘要 · Abstract (English)
High-fidelity 3D reconstruction is critical for aerial inspection tasks such as infrastructure monitoring, structural assessment, and environmental surveying. While traditional photogrammetry techniques enable geometric modeling, they lack semantic interpretability, limiting their effectiveness for automated inspection workflows. Recent advances in neural rendering and 3D Gaussian Splatting (3DGS) offer efficient, photorealistic reconstructions but similarly lack scene-level understanding. In this work, we present a UAV-based pipeline that extends Feature-3DGS for language-guided 3D segmentation. We leverage LSeg-based feature fields with CLIP embeddings to generate heatmaps in response to language prompts. These are thresholded to produce rough segmentations, and the highest-scoring point is then used as a prompt to SAM or SAM2 for refined 2D segmentation on novel view renderings. Our results highlight the strengths and limitations of various feature field backbones (CLIP-LSeg, SAM, SAM2) in capturing meaningful structure in large-scale outdoor environments. We demonstrate that this hybrid approach enables flexible, language-driven interaction with photorealistic 3D reconstructions, opening new possibilities for semantic aerial inspection and scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。