用大规模数据训练通用模型,无需特殊架构即可在遥感任务上达到顶尖表现。
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

- 用统一语言策略和多任务强化学习训练,支持问答与定位双重行为。
- 在高分辨率、多时相等多类遥感任务中表现媲美或超越现有方法。
- 数据规模比模型架构更重要,适合追求高效实用的遥感研究者。
遥感视觉-语言模型需支持对地球观测数据的开放域推理与多样化任务。当前进展多依赖专用架构设计,如新编码器、对齐模块或任务特异性融合机制。本文挑战这一必要性:仅通过在多样数据与任务上大规模训练,通用视觉-语言模型即可在挑战性遥感基准上达到竞争性或领先性能。模型采用单一语言策略,可直接输出文本答案或调用定位工具完成分割与定位。为训练这种异构行为,我们提出基于自适应任务奖励的多任务强化学习框架,覆盖多项选择式VQA、自由问答、描述生成、检测与分割,涵盖多种输入类型。实验表明,该方法在广泛基准上表现优异,包括高分辨率、多时相、多模态与多视角任务。随着训练数据规模增加,多数任务(包括分布外)性能持续提升,且与每项任务的数据多样性正相关。结果表明,对遥感视觉-语言模型而言,数据规模的重要性超过架构创新。
原文摘要 · Abstract (English)
Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can achieve competitive or state-of-the-art performance at challenging remote sensing benchmarks, provided that it is trained at sufficient scale across diverse data and tasks. Our model uses a single language policy that can either answer directly in text or invoke a localization tool for segmentation and grounding. To train this heterogeneous behaviour, we employ a multi-task reinforcement learning framework with adaptive task rewards covering multiple-choice VQA, free-form VQA, captioning, detection, and segmentation across a large variety of input types. Our approach achieves competitive results across a broad set of benchmarks, including high-resolution, multi-temporal, multi-modal and multi-view tasks. Further, as training data scales, our experiments show consistent improvements across most tasks both in and out of distribution, which correlate with per-task data diversity. These findings suggest that, for remote sensing VLMs, data scale is more important than architectural novelty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。