首个跨域遥感视觉定位框架,解决光学与雷达图像匹配难题
OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding
- 用低秩混合专家模块实现跨域特征解耦,提升模型效率
- 在两个基准数据集上达到当前最优性能,定位准确率显著提升
- 适合遥感智能解译、多模态融合研究者参考
遥感视觉定位(RSVG)旨在通过自然语言描述定位遥感图像中的目标。现有方法局限于单一传感器域(光学或合成孔径雷达),限制了实际应用。本文提出跨域遥感视觉定位(CD-RSVG)任务,并构建首个大规模基准数据集OptSAR-RSVG。为应对跨域特征建模、计算效率和细粒度语义区分挑战,提出OptiSAR-Net++框架。该框架采用分块级低秩自适应混合专家(PL-MoE)实现高效跨域特征解耦;摒弃Transformer解码结构,采用基于CLIP的对比学习范式并引入动态对抗负采样,将生成回归转化为高效的跨模态匹配;还设计文本引导双门控融合模块(TGDF-SSA)与区域感知辅助头,增强语义-视觉对齐与空间建模。大量实验表明,OptiSAR-Net++在OptSAR-RSVG与DIOR-RSVG两个基准上均达当前最优表现,定位精度与效率均有显著提升。代码与数据集将公开。
原文摘要 · Abstract (English)
Remote sensing visual grounding (RSVG) aims to localize specific targets in remote sensing images using natural language expressions. However, existing methods are restricted to single-sensor domains, i.e., either optical or synthetic aperture radar (SAR), limiting their real-world applicability. In this paper, we introduce the Cross-Domain RSVG (CD-RSVG) task and construct OptSAR-RSVG, the first large-scale benchmark dataset for this setting. To tackle the challenges of cross-domain feature modeling, computational inefficiency, and fine-grained semantic discrimination, we propose OptiSAR-Net++. Our framework features a patch-level Low-Rank Adaptation Mixture of Experts (PL-MoE) for efficient cross-domain feature decoupling. To mitigate the substantial computational overhead of Transformer decoding frameworks, we adopt a CLIP-based contrastive paradigm and further incorporate dynamic adversarial negative sampling, thereby transforming generative regression into an efficient cross-modal matching process. Additionally, a text-guided dual-gate fusion module (TGDF-SSA) and a region-aware auxiliary head are introduced to enhance semantic-visual alignment and spatial modeling. Extensive experiments demonstrate that OptiSAR-Net++ achieves SOTA performance on both OptSAR-RSVG and DIOR-RSVG benchmarks, offering significant advantages in localization accuracy and efficiency. Our code and dataset will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。