arXiv:2505.12099cs.CV2025-05中稿 · IEEE Geoscience an…被引 6

20亿参数小模型,专为遥感任务优化,性能媲美70亿参数大模型。

TinyRS-R1: Compact Multimodal Language Model for Remote Sensing

  • 基于Qwen2-VL-2B构建,四阶段训练提升遥感理解能力。
  • 在分类、问答等任务上超越7B模型,内存和延迟仅需三分之一。
  • 适合边缘设备部署,尤其擅长空间定位与复杂场景理解。

遥感应用常运行在无法支持当前70亿参数多模态语言模型的边缘硬件上。本文提出首个20亿参数的遥感专用小规模多模态语言模型TinyRS及其推理增强版TinyRS-R1。TinyRS基于Qwen2-VL-2B,采用四阶段训练流程:在百万张卫星图像上预训练,使用视觉指令样本进行指令微调,利用自建推理数据集中的思维链(CoT)标注进行微调,并通过分组相对策略优化(GRPO)对齐。TinyRS-R1在分类、视觉问答(VQA)、视觉定位和开放问答任务上达到或超越近期70亿参数遥感模型的性能,同时仅需其三分之一的内存与延迟。分析表明,CoT推理显著提升空间定位与场景理解能力,而无推理版本的TinyRS则更适用于简洁、低延迟的VQA任务。TinyRS-R1是首个经GRPO对齐的具有思维链推理能力的通用遥感领域专用小模型。

原文摘要 · Abstract (English)

Remote-sensing applications often run on edge hardware that cannot host today's 7B-parameter multimodal language models. This paper introduces TinyRS, the first 2B-parameter multimodal small language model (MSLM) optimized for remote sensing tasks, and TinyRS-R1, its reasoning-augmented variant. Built upon Qwen2-VL-2B, TinyRS is trained through a four-stage pipeline: pre-training on million satellite images, instruction tuning on visual instruction examples, fine-tuning with Chain-of-Thought (CoT) annotations from the proposed reasoning dataset, and alignment via Group Relative Policy Optimization (GRPO). TinyRS-R1 achieves or surpasses the performance of recent 7B-parameter remote sensing models across classification, VQA, visual grounding, and open-ended question answering-while requiring just one-third of the memory and latency. Our analysis shows that CoT reasoning substantially benefits spatial grounding and scene understanding, while the non-reasoning TinyRS excels in concise, latency-sensitive VQA tasks. TinyRS-R1 represents the first domain-specialized MSLM with GRPO-aligned CoT reasoning for general-purpose remote sensing.

遥感小模型多模态推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。