arXiv:2607.16819cs.AI2026-07

FUSAR-R1让SAR图像解读更像人类思考,提升复杂场景下识别可靠性。

FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images

论文配图:FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images
图 1 · 摘自论文原文
  • 模拟专家推理过程构建思维链数据,训练模型分步分析能力。
  • 引入强化学习优化推理结果,实现自我修正与更可靠判断。
  • 适合需要高精度、可解释性SAR图像分析的科研与应用人群。

近年来,大规模视觉-语言模型正推动智能遥感图像解译范式变革。通过融合文本语义信息,解译模型的认知表达、语义理解与人机交互能力显著提升,已在合成孔径雷达(SAR)图像解译领域取得初步进展。然而,SAR图像受相干成像机制、复杂散射特性、斑点噪声干扰及目标-背景耦合等因素影响,特征复杂多变,存在显著不确定性与专业化特点。现有SAR视觉-语言模型尚不具备人类专家的逐步分析、逻辑判断与自我纠错能力,难以在复杂场景中支撑可靠智能解译。为此,本文提出面向SAR图像智能解译的大规模推理模型FUSAR-R1。该模型首先通过模拟人类专家的解译流程构建显式思维链(chain-of-thought)推理数据,并利用该数据指导指令学习,赋予模型基础推理能力;随后引入强化学习策略,基于推理结果优化模型输出,实现自我修正与更可靠的推理。实验表明,FUSAR-R1在目标检测、目标计数与分类、地表覆盖类别识别等多种SAR解译任务上,持续优于现有多模态大模型。

原文摘要 · Abstract (English)

In recent years, large-scale vision-language models have been driving a paradigm shift in intelligent remote sensing image interpretation. By incorporating textual semantic information, the cognitive expression, semantic understanding, and human-computer interaction capabilities of interpretation models have been significantly improved, achieving initial progress in the field of Synthetic Aperture Radar (SAR) image interpretation. However, SAR images are affected by factors such as coherent imaging mechanisms, complex scattering characteristics, speckle noise interference, and target-background coupling, resulting in complex and variable image features with significant uncertainties and specializations. Existing SAR vision-language models do not yet possess the step-by-step analysis, logical judgment, and self-correction capabilities of human experts, making it difficult to support reliable intelligent interpretation in complex scenarios. To address this issue, this paper proposes a large-scale reasoning model, FUSAR-R1, for intelligent interpretation of SAR images. The model first constructs explicit chain-of-thought reasoning data by simulating the interpretation process of human experts and uses this data to guide instruction learning, thereby endowing the model with basic reasoning capabilities. Subsequently, a reinforcement learning strategy is introduced to optimize the model's outputs based on inference results, enabling self-correction and more reliable reasoning. Experimental results demonstrate that FUSAR-R1 consistently outperforms existing multimodal large-scale models across various SAR interpretation tasks, including target detection, target counting and classification, and land-cover category recognition.

SAR图像推理模型视觉语言强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。