arXiv:2603.01758cs.CV2026-03

用语言做桥梁,统一多源遥感目标检测

Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining

论文配图:Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
图 1 · 摘自论文原文
  • 以语言为语义枢纽,解耦模态对齐与任务学习
  • 在多个数据集上超越现有方法,训练更稳定
  • 适合多传感器遥感检测研究者参考

异构多模态遥感目标检测旨在从不同传感器(如可见光、SAR、红外)中准确识别物体。现有方法多采用晚期对齐范式,导致模态对齐与下游任务优化在微调阶段纠缠,引发训练不稳定和泛化性能不佳。为此,我们提出BabelRS,一种统一的语言引导预训练框架,显式解耦模态对齐与下游任务学习。BabelRS包含两个核心组件:概念共享指令对齐(CSIA)和逐层视觉-语义退火(LVSA)。CSIA通过语言作为语义枢纽,将各传感器模态对齐至共享的语义概念空间;为缓解高层语言表示与密集检测目标间的粒度差异,LVSA逐步聚合多尺度视觉特征,提供细粒度语义引导。大量实验表明,BabelRS显著提升训练稳定性,且在无额外技巧的情况下持续优于当前最优方法。代码已开源:https://github.com/zcablii/SM3Det。

原文摘要 · Abstract (English)

Heterogeneous multi-modal remote sensing object detection aims to accurately detect objects from diverse sensors (e.g., RGB, SAR, Infrared). Existing approaches largely adopt a late alignment paradigm, in which modality alignment and task-specific optimization are entangled during downstream fine-tuning. This tight coupling complicates optimization and often results in unstable training and suboptimal generalization. To address these limitations, we propose BabelRS, a unified language-pivoted pretraining framework that explicitly decouples modality alignment from downstream task learning. BabelRS comprises two key components: Concept-Shared Instruction Aligning (CSIA) and Layerwise Visual-Semantic Annealing (LVSA). CSIA aligns each sensor modality to a shared set of linguistic concepts, using language as a semantic pivot to bridge heterogeneous visual representations. To further mitigate the granularity mismatch between high-level language representations and dense detection objectives, LVSA progressively aggregates multi-scale visual features to provide fine-grained semantic guidance. Extensive experiments demonstrate that BabelRS stabilizes training and consistently outperforms state-of-the-art methods without bells and whistles. Code: https://github.com/zcablii/SM3Det.

遥感检测多模态语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。