arXiv:2501.14905cs.CV2025-01被引 3

用地图增强遥感图文数据生成,减少大模型幻觉。

Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing

  • 引入地图作为外部信息源,生成更精准的遥感描述
  • 提出评估与抑制大模型幻觉的方法,提升文本真实性
  • 新数据集fMoW-mm在少样本目标识别中表现更优

视觉语言模型在多个领域取得显著成果,但在遥感领域的应用受限于图像-文本配对数据稀缺。为弥补这一缺口,合成文本生成受到关注,传统方法依赖规则,使用元数据或边界框生成描述,但难以捕捉复杂大范围场景。大语言模型(LLMs)可生成更丰富描述,却易产生通用化输出和幻觉。本文提出一种新方法,通过整合地图作为外部数据源,生成细节丰富、上下文准确的遥感图像描述。同时,设计了衡量与缓解LLM生成文本幻觉的策略。构建了fMoW-mm多模态数据集,包含卫星影像、地图、元数据及文本标注。实验表明,该数据集在少样本自动目标识别任务中优于现有遥感视觉语言数据集。

原文摘要 · Abstract (English)

Vision language models have achieved impressive results across various fields. However, adoption in remote sensing remains limited, largely due to the scarcity of paired image-text data. To bridge this gap, synthetic caption generation has gained interest, traditionally relying on rule-based methods that use metadata or bounding boxes. While these approaches provide some description, they often lack the depth needed to capture complex wide-area scenes. Large language models (LLMs) offer a promising alternative for generating more descriptive captions, yet they can produce generic outputs and are prone to hallucination. In this paper, we propose a new method to enhance vision-language datasets for remote sensing by integrating maps as external data sources, enabling the generation of detailed, context-rich captions. Additionally, we present methods to measure and mitigate hallucinations in LLM-generated text. We introduce fMoW-mm, a multimodal dataset incorporating satellite imagery, maps, metadata, and text annotations. We demonstrate its effectiveness for automatic target recognition in few-shot settings, achieving superior performance compared to other vision-language remote sensing datasets.

遥感大模型幻觉抑制图文生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。