用多模型协作提升遥感图像语义定位精度
Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting

- 融合RemoteSAM与SAM3,分步优化遥感目标定位
- 六模型集成投票使定位准确率显著提升
- 适合需要高精度遥感语义理解的科研与应用
视觉定位旨在找出与自然语言描述对应图像区域,是可解释视觉系统的关键。在遥感图像中,由于场景复杂、目标小且尺度变化大,定位尤为困难。单一模型难以应对这些挑战。本文提出两种新框架:顺序精炼(SGR)与聚类感知精炼(CGR),结合专用于遥感的Visual Grounding模型RemoteSAM和通用分割模型SAM3的互补优势。首先用RemoteSAM获取目标初始位置,再通过SAM3进行精细化修正,得到更精确、空间一致的分割结果。此外,我们设计基于多数投票的集成策略,融合六个具有不同能力的定位管道。该多模型框架增强了鲁棒性,显著提升定位准确性。实验表明,所提方法优于单个模型,实现更可靠、精准的视觉定位预测。
原文摘要 · Abstract (English)
Visual grounding aims to locate image regions that correspond to natural language descriptions and is a key component of interpretable vision systems. In remote sensing imagery, grounding is particularly challenging due to complex scenes, small objects, and large variations in scale. Relying on a single model is often insufficient to address these diverse challenges. In this work, we propose two grounding pipelines, Sequential Grounding Refinement (SGR) and Cluster-Aware Grounding Refinement (CGR), that combine the complementary strengths of RemoteSAM, a visual grounding model specialized for remote sensing, and SAM3, a powerful general-purpose segmentation model. Our approach first uses RemoteSAM to obtain an initial estimate of object location, which is then refined using SAM3 to produce more accurate and spatially consistent segmentations. Additionally, we explore an ensemble strategy based on majority voting across six diverse grounding pipelines, each with distinct capabilities. This multi-model framework improves robustness and significantly enhances localization accuracy. Experimental results demonstrate that the proposed pipelines and ensemble approach outperform individual models, leading to more reliable and precise visual grounding predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。