arXiv:2608.25435cs.CVcs.AI2026-08

用视觉显著性与深度信息提升无人机图像中通信塔部件的零样本分割准确率

Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery

论文配图:Saliency-Depth Conditioning for Zero-Shot Segmentation of Communication-Tower Components in Cluttered UAV Imagery
图 1 · 摘自论文原文
  • 融合显著性与单目深度图生成塔体粗略先验,抑制杂乱背景干扰
  • 在TOW-300数据集上,新模型实例分割精度提升,误检率显著降低
  • 适合缺乏标注数据的通信塔自动化巡检场景,尤其适用于复杂背景

无人机图像中精细分割通信塔部件对自动化巡检至关重要,但因实例级标注稀缺,定制模型难训练。零样本分割模型虽具潜力,但在杂乱场景中,外观相似的背景结构会干扰部件定位,导致漏检和误检。本文提出一种模型无关的显著性-深度前景条件策略,结合基于外观的显著性与单目相对深度,构建粗粒度塔体先验,有效抑制无关内容。将该模块集成至Grounded-SAM与SAM 3,得到SD-Grounded-SAM与SD-SAM 3。SD-Grounded-SAM在生成掩码前引入几何与深度感知的框精修,而SD-SAM 3依赖SAM 3内部机制。在包含340张通信塔无人机图像的TOW-300数据集上,所提方法均优于基线:SD-SAM 3达到最强实例分割性能,而SD-Grounded-SAM误检更少。消融实验验证了显著性、深度与框精修的互补增益,显著提升了复杂场景下的鲁棒性。

原文摘要 · Abstract (English)

Fine-grained segmentation of communication-tower components in UAV imagery is essential for automated inspection, yet task-specific models are hard to develop due to limited instance-level annotations. Zero-shot segmentation models offer a promising alternative, but in cluttered scenes, visually similar background structures interfere with component localization, causing missed instances and false positives. We propose a model-agnostic saliency-depth foreground-conditioning strategy combining appearance-based saliency with monocular relative depth to construct a coarse tower prior and suppress irrelevant content. We integrate this module with Grounded-SAM and SAM 3, yielding SD-Grounded-SAM and SD-SAM 3. SD-Grounded-SAM further applies geometric and depth-aware box refinement before mask generation, while SD-SAM 3 relies on SAM 3's internal setup. On TOW-300, a dataset of 340 communication-tower UAV images, our strategy improves both baselines: SD-SAM 3 achieves the strongest instance-segmentation performance, while SD-Grounded-SAM produces fewer false positives. Ablations confirm complementary gains from saliency, depth, and box refinement, improving robustness in cluttered scenes.

零样本分割无人机图像通信塔检测深度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。