arXiv:2505.07984cs.CV2025-05中稿 · Journal of Selecte…被引 5

轻量多模态模型通过思维链与强化学习,精准识别偏远地区军事设施。

SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing

  • 引入思维链提示与GRPO强化学习,提升模型推理能力。
  • 在新数据集上实现80%召回率、98%精确率,显著优于大模型。
  • 适合需要低资源、高精度的遥感图像分析场景。

近期多模态大模型在图文理解与生成方面表现卓越,但在资源受限且需领域特化的场景中效果有限。本文提出轻量级多模态语言模型SAMChat,专用于分析偏远地区遥感影像,包括复杂导弹发射场。构建了新数据集SAMData,经专家审核数百张航拍图,标注细微军事设施并附详细描述。在20亿参数开源多模态模型上进行带思维链(CoT)注释的监督微调,增强解释性与准确性。同时采用分组相对策略优化(GRPO),提升对防御布局、关键军事结构等特定线索的检测能力,减少民用场景误报。实证表明,SAMChat在开放性描述与分类任务上显著优于更大通用模型及现有遥感适配方法,在新基准SAMData上达到80%召回率与98%精确率,验证了针对性微调与强化学习在真实世界应用中的有效性。

原文摘要 · Abstract (English)

Remarkable capabilities in understanding and generating text-image content have been demonstrated by recent advancements in multimodal large language models (MLLMs). However, their effectiveness in specialized domains-particularly those requiring resource-efficient and domain-specific adaptations-has remained limited. In this work, a lightweight multimodal language model termed SAMChat is introduced, specifically adapted to analyze remote sensing imagery in secluded areas, including challenging missile launch sites. A new dataset, SAMData, was compiled by verifying hundreds of aerial images through expert review, and subtle military installations were highlighted via detailed captions. Supervised fine-tuning on a 2B parameter open-source MLLM with chain-of-thought (CoT) reasoning annotations was performed, enabling more accurate and interpretable explanations. Additionally, Group Relative Policy Optimization (GRPO) was leveraged to enhance the model's ability to detect critical domain-specific cues-such as defensive layouts and key military structures-while minimizing false positives on civilian scenes. Through empirical evaluations, it has been shown that SAMChat significantly outperforms both larger, general-purpose multimodal models and existing remote sensing adapted approaches on open-ended captioning and classification metrics. Over 80% recall and 98% precision were achieved on the newly proposed SAMData benchmark, underscoring the potency of targeted fine-tuning and reinforcement learning in specialized real-world applications.

遥感分析多模态模型强化学习轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。