用学习的评分模型自动筛选高质量遥感图文数据,提升视觉语言模型性能。
Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models
- 基于遥感图文偏好数据训练评分模型,自动评估合成数据质量。
- 用评分模型选前30%数据微调,精度优于全量数据或传统评分方法。
- 可应用于强化学习与测试阶段提升,适合遥感多模态研究者使用。
视觉语言模型(VLM)在语义引导下展现了解析遥感(RS)图像的巨大潜力。然而,其性能高度依赖于高质量的图文训练数据,该数据需准确捕捉视觉内容与语言描述间的丰富语义关系。与自然图像不同,遥感领域缺乏来自网络的大规模交错图文对,导致数据收集困难。现有方法主要依赖规则或主流VLM进行数据合成,但缺乏系统性的自动化质量评估框架。为此,我们提出一种基于大规模遥感图文偏好数据训练的评分模型,实现对合成遥感图文数据的质量自动评估。实验表明,使用本模型排序后前30%的数据微调CLIP或先进VLM(如Qwen2-VL),性能优于全量数据微调及基于CLIP-score的排名方法。此外,我们还展示了该评分模型在强化学习训练和最佳- N(BoN)测试时缩放中的应用,显著提升了遥感任务中VLM的表现。代码、模型与数据集均已公开。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have demonstrated great potential in interpreting remote sensing (RS) images through language-guided semantic. However, the effectiveness of these VLMs critically depends on high-quality image-text training data that captures rich semantic relationships between visual content and language descriptions. Unlike natural images, RS lacks large-scale interleaved image-text pairs from web data, making data collection challenging. While current approaches rely primarily on rule-based methods or flagship VLMs for data synthesis, a systematic framework for automated quality assessment of such synthetically generated RS vision-language data is notably absent. To fill this gap, we propose a novel score model trained on large-scale RS vision-language preference data for automated quality assessment. Our empirical results demonstrate that fine-tuning CLIP or advanced VLMs (e.g., Qwen2-VL) with the top 30% of data ranked by our score model achieves superior accuracy compared to both full-data fine-tuning and CLIP-score-based ranking approaches. Furthermore, we demonstrate applications of our scoring model for reinforcement learning (RL) training and best-of-N (BoN) test-time scaling, enabling significant improvements in VLM performance for RS tasks. Our code, model, and dataset are publicly available
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。