先学地球科学文本再训练,让模型更准理解超高清遥感图像。
Text Before Vision: Staged Knowledge Injection Matters for Agentic RLVR in Ultra-High-Resolution Remote Sensing Understanding
- 先用高质量地球科学文本预训练,建立推理框架
- 在超高清遥感数据上实现60.40%的通过率,领先大模型
- 适合遥感、地理信息等需要复杂视觉推理的领域
超高清遥感(UHR RS)多模态推理常受限于视觉证据获取:模型需在海量像素中定位微小任务相关区域。尽管使用可验证奖励的代理强化学习(RLVR)和缩放工具提供了新路径,我们发现标准强化学习在缺乏结构化领域先验时难以导航广阔视觉空间。本研究对比了冷启动监督微调(SFT)、RLVR与代理式RLVR在UHR RS基准上的表现。控制实验揭示反直觉发现:仅含文本的地球科学问答数据是提升超高清视觉推理能力的关键驱动力。即便无图像,特定领域的文本仍能注入概念、机制解释与决策规则,引导视觉证据检索。基于此,我们提出分阶段知识注入方案:(1) 利用可扩展、知识图谱验证的地球科学文本问答进行冷启动,建立推理结构;(2) 在相同困难的超高清图文样本上进行微调以稳定并放大后续工具型强化学习效果。该方法在XLRS-Bench上达到60.40%的Pass@1,显著优于更大通用模型(如GPT-5.2、Gemini 3.0 Pro、Intern-S1),刷新当前最优水平。
原文摘要 · Abstract (English)
Multimodal reasoning for ultra-high-resolution (UHR) remote sensing (RS) is usually bottlenecked by visual evidence acquisition: the model necessitates localizing tiny task-relevant regions in massive pixel spaces. While Agentic Reinforcement Learning with Verifiable Rewards (RLVR) using zoom-in tools offers a path forward, we find that standard reinforcement learning struggles to navigate these vast visual spaces without structured domain priors. In this paper, we investigate the interplay between post-training paradigms: comparing Cold-start Supervised Fine-Tuning (SFT), RLVR, and Agentic RLVR on the UHR RS benchmark.Our controlled studies yield a counter-intuitive finding: high-quality Earth-science text-only QA is a primary driver of UHR visual reasoning gains. Despite lacking images, domain-specific text injects the concepts, mechanistic explanations, and decision rules necessary to guide visual evidence retrieval.Based on this, we propose a staged knowledge injection recipe: (1) cold-starting with scalable, knowledge-graph-verified Earth-science text QA to instill reasoning structures;and (2) "pre-warming" on the same hard UHR image-text examples during SFT to stabilize and amplify subsequent tool-based RL. This approach achieves a 60.40% Pass@1 on XLRS-Bench, significantly outperforming larger general purpose models (e.g., GPT-5.2, Gemini 3.0 Pro, Intern-S1) and establishing a new state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。