arXiv:2607.22205cs.CVcs.AI2026-07

针对遥感模型场景专精难题,先补能力短板再训练,效果更优。

Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs

论文配图:Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs
图 1 · 摘自论文原文
  • 分三阶段逐步填补模型能力缺口,再进行场景化训练。
  • 在港口遥感任务中,性能从57.95提升至70.29(LLaVA-v1.5)。
  • 适合需要高精度遥感理解的科研与工程应用者。

遥感多模态大模型虽提升了通用空域图像理解能力,但地球观测需细粒度场景专精,受限于高质量场景数据稀缺和能力覆盖不全。本文将适配问题建模为能力缺口驱动的后训练任务,提出“先填后进”(FBA)方法:不依赖单一阶段监督微调,而是先填补关键能力缺口,再推进场景专精。以典型多源港口场景为例,构建三层监督数据集CPRS,包含三个有序阶段:(1) 遥感语义锚定,实现俯视视觉-语言对齐;(2) 域桥收敛,跨模态对齐目标与桥接场景的共享遥感先验;(3) 证据驱动的场景调优,提升下游表现。构建八项指标的HarborEval诊断基准,涵盖感知、空间理解、鲁棒性与生成能力。在相近训练预算下,相较于直接微调(Direct-SFT),FBA使LLaVA-v1.5在HarborEval上从57.95提升至70.29,Qwen3-VL从81.09升至83.37。FBA优于崩溃式微调(Collapsed-SFT),并在harbor相关VRSBench/RSVQA子集及OpenEval上领先。阶段分析与角色替换验证了渐进式补缺与阶段特异性作用。CPRS、HarborEval、代码与模型权重已开源。

原文摘要 · Abstract (English)

Remote sensing multimodal large language models (RS-MLLMs) have improved general aerial-image understanding. However, Earth observation applications require fine-grained scenario specialization, constrained by scarce high-quality scenario data and incomplete capability coverage. We formulate this adaptation as a capability-gap-driven post-training problem and propose filling before advancing (FBA). Rather than relying on single-stage supervised fine-tuning (SFT) over target-domain samples, FBA first fills prerequisite capability gaps before advancing toward scenario specialization. We instantiate FBA for coastal harbor understanding, a representative multi-source scenario, by constructing CPRS (Coastal-Port Remote Sensing), a three-layer supervision dataset coupled with three ordered stages: (1) RS semantic anchoring for overhead-view visual-language alignment; (2) domain-bridge convergence for shared RS priors across target and bridging scenarios under different modalities; and (3) evidence-grounded scenario tuning for downstream performance. We construct HarborEval, an eight-track diagnostic benchmark covering perception, spatial understanding, robustness, and generation. Under comparable training budgets, HarborEval increases from 57.95 with Direct-SFT to 70.29 with FBA on LLaVA-v1.5, and from 81.09 to 83.37 on Qwen3-VL. FBA also outperforms Collapsed-SFT and leads on harbor-related VRSBench/RSVQA subsets and OpenEval. Stage-wise and role-replacement analyses validate progressive gap filling and stage-specific roles. Public examples and release updates for CPRS, HarborEval, code, and trained weights are available at https://github.com/Z0ngL1ng/filling-before-advancing.

遥感多模态后训练场景专精

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。