用少量样本+可验证奖励,让卫星图像模型学会推理。
Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
- 仅需1个标注样例,通过可验证奖励训练视觉语言模型
- 1个样例提升显著,128个样例媲美数千样本训练效果
- 适合数据稀缺的遥感领域,降低标注成本
大型语言与视觉语言模型虽具强大推理能力,但在遥感等专业领域因标注数据稀缺而难应用。本文提出首个面向卫星影像的少样本可验证奖励强化学习(RLVR)框架,无需图像描述监督,仅依赖轻量级规则化二值或IoU奖励。将语言模型中的“1样本RLVR”范式拓展至视觉语言模型,利用策略梯度优化,仅需一个精心挑选的示例即可对齐模型输出以完成卫星影像推理任务。在多个遥感基准测试(包括分类、视觉问答和定位)中,即使单样本也显著优于基础模型;128样本性能达到甚至超过数千标注样本训练的模型。尽管极少数样本可能引发轻微任务特定过拟合,但整体展现稳健泛化能力与高效性。此外,提示设计与损失权重对训练稳定性和最终精度影响显著。本方法为数据稀缺领域提供低成本、高效率的专用视觉语言推理模型开发方案:从紧凑的VLM出发,精选少量可验证奖励案例,通过RLVR训练。
原文摘要 · Abstract (English)
Recent advances in large language and vision-language models have enabled strong reasoning capabilities, yet they remain impractical for specialized domains like remote sensing, where annotated data is scarce and expensive. We present the first few-shot reinforcement learning with verifiable reward (RLVR) framework for satellite imagery that eliminates the need for caption supervision--relying solely on lightweight, rule-based binary or IoU-based rewards. Adapting the "1-shot RLVR" paradigm from language models to vision-language models, we employ policy-gradient optimization with as few as one curated example to align model outputs for satellite reasoning tasks. Comprehensive experiments across multiple remote sensing benchmarks--including classification, visual question answering, and grounding--show that even a single example yields substantial improvements over the base model. Scaling to 128 examples matches or exceeds models trained on thousands of annotated samples. While the extreme one-shot setting can induce mild, task-specific overfitting, our approach consistently demonstrates robust generalization and efficiency across diverse tasks. Further, we find that prompt design and loss weighting significantly influence training stability and final accuracy. Our method enables cost-effective and data-efficient development of domain-specialist vision-language reasoning models, offering a pragmatic recipe for data-scarce fields: start from a compact VLM, curate a handful of reward-checkable cases, and train via RLVR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。