构建首个中英双语遥感图像描述数据集,推动多语言遥感理解发展。
A Benchmark for Multi-Lingual Vision-Language Learning in Remote Sensing Image Captioning
- 构建包含13,634张图像的中英双语遥感图像描述数据集
- 在8个主流大模型上验证多语言能力,发现跨语言迁移存在显著差异
- 提供统一评估框架,适合关注多语言遥感应用的研究者
遥感图像描述(RSIC)是视觉与语言交叉领域,旨在自动生成遥感图像中的场景和特征的自然语言描述。尽管在开发复杂方法和大规模数据集方面取得进展,但非英语描述数据稀缺、模型多语言能力缺乏评估仍是两大挑战,严重制约了RSIC的发展与实际应用。本文提出BRSIC(双语遥感图像描述)数据集,将三个现有英文数据集扩展为中英双语,涵盖13,634张图像及68,170条双语描述。基于此,建立系统化评估框架,通过标准化重训练流程实现严格模型性能评测。进一步对8个前沿大视觉语言模型(LVLMs)进行广泛实证研究,涵盖零样本推理、监督微调与多语言训练等多种范式,揭示当前模型在多语言遥感任务中的优劣势。跨数据集迁移实验也获得重要发现。代码与数据将在https://github.com/mrazhou/BRSIC公开。
原文摘要 · Abstract (English)
Remote Sensing Image Captioning (RSIC) is a cross-modal field bridging vision and language, aimed at automatically generating natural language descriptions of features and scenes in remote sensing imagery. Despite significant advances in developing sophisticated methods and large-scale datasets for training vision-language models (VLMs), two critical challenges persist: the scarcity of non-English descriptive datasets and the lack of multilingual capability evaluation for models. These limitations fundamentally impede the progress and practical deployment of RSIC, particularly in the era of large VLMs. To address these challenges, this paper presents several significant contributions to the field. First, we introduce and analyze BRSIC (Bilingual Remote Sensing Image Captioning), a comprehensive bilingual dataset that enriches three established English RSIC datasets with Chinese descriptions, encompassing 13,634 images paired with 68,170 bilingual captions. Building upon this foundation, we develop a systematic evaluation framework that addresses the prevalent inconsistency in evaluation protocols, enabling rigorous assessment of model performance through standardized retraining procedures on BRSIC. Furthermore, we present an extensive empirical study of eight state-of-the-art large vision-language models (LVLMs), examining their capabilities across multiple paradigms including zero-shot inference, supervised fine-tuning, and multi-lingual training. This comprehensive evaluation provides crucial insights into the strengths and limitations of current LVLMs in handling multilingual remote sensing tasks. Additionally, our cross-dataset transfer experiments reveal interesting findings. The code and data will be available at https://github.com/mrazhou/BRSIC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。