针对遥感多模态大模型幻觉问题,提出系统评估与缓解方案。
RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation
- 构建遥感专属幻觉分类体系,涵盖图像级语义不一致
- 发布2023对问答数据集,支持高精度云检测与本地低成本检查
- 提出训练免费的幻觉缓解策略,适配遥感场景应用
多模态大语言模型在遥感领域应用日益广泛,但在应急管理和农业监测等高风险场景中,因输入图像与输出响应不符导致的幻觉问题严重制约其部署,且该问题在遥感领域仍缺乏深入研究。本文提出RSHallu,包含三项成果:(1) 构建面向遥感的幻觉分类体系,引入图像级幻觉以捕捉超越目标中心错误的模态、分辨率和场景级语义不一致;(2) 构建包含2,023组问答对的幻觉评估基准RSHalluEval,支持双模式检测——通过在15,396组问答对上微调的紧凑检查器实现高精度云端审计,同时支持低成本可复现的本地检查;(3) 提出面向训练友好的领域定制数据集RSHalluShield(30k问答对),并设计无需训练的即插即用策略,包括解码时的逻辑纠正和遥感感知提示。在多个代表性遥感多模态大模型上,该缓解方法在统一协议下将无幻觉率提升最高达21.63个百分点,同时保持下游任务(RSVQA/RSVG)的竞争力。代码与数据集将公开。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are increasingly adopted in remote sensing (RS) and have shown strong performance on tasks such as RS visual grounding (RSVG), RS visual question answering (RSVQA), and multimodal dialogue. However, hallucinations, which are responses inconsistent with the input RS images, severely hinder their deployment in high-stakes scenarios (e.g., emergency management and agricultural monitoring) and remain under-explored in RS. In this work, we present RSHallu, a systematic study with three deliverables: (1) we formalize RS hallucinations with an RS-oriented taxonomy and introduce image-level hallucination to capture RS-specific inconsistencies beyond object-centric errors (e.g., modality, resolution, and scene-level semantics); (2) we build a hallucination benchmark RSHalluEval (2,023 QA pairs) and enable dual-mode checking, supporting high-precision cloud auditing and low-cost reproducible local checking via a compact checker fine-tuned on RSHalluCheck dataset (15,396 QA pairs); and (3) we introduce a domain-tailored dataset RSHalluShield (30k QA pairs) for training-friendly mitigation and further propose training-free plug-and-play strategies, including decoding-time logit correction and RS-aware prompting. Across representative RS-MLLMs, our mitigation improves the hallucination-free rate by up to 21.63 percentage points under a unified protocol, while maintaining competitive performance on downstream RS tasks (RSVQA/RSVG). Code and datasets will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。