arXiv:2606.20177cs.CVcs.AI2026-06中稿 · ed

首个遥感多模态模型否定理解评测基准,提升模型对不存在信息的识别能力。

Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs

论文配图:Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
图 1 · 摘自论文原文
  • 构建自动化数据生成管道,合成多样化否定查询并动态验证视觉焦点。
  • 先进模型在否定任务上性能下降明显,存在严重幻觉现象。
  • 提出NeFo方法,仅用5%无标签样本即可显著提升否定理解与泛化能力。

多模态大语言模型(MLLMs)在遥感(RS)任务中表现卓越,但其否定理解能力尚未得到充分研究,限制了在真实场景中的应用——例如应急响应人员需识别未被洪水淹没的道路。为此,我们提出首个涵盖区域级至场景级任务的否定理解评测基准RS-Neg。通过使用大语言模型自动生成遥感图像中的多样化否定查询,并引入动态视觉聚焦模块进行验证。评估显示,先进的遥感MLLMs在否定任务中表现不佳,存在显著幻觉和性能下降。为弥补这一差距,我们提出NeFo,一种测试时学习方法,显式将否定逻辑融入模型优化。令人惊讶的是,仅需约5%的无标签测试样本,NeFo即能显著提升模型的否定理解能力,并在未见任务上表现出强泛化性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability to comprehend negation remains underexplored, limiting deployment in real-world applications where models must explicitly identify what is false or absent, e.g., emergency responders need to locate non-flooded routes for evacuation. To comprehensively study this limitation, we introduce RS-Neg, the first benchmark to evaluate negation understanding across region-level to scene-level tasks. Specifically, we design an automated data generation pipeline for RS imagery, using LLMs to synthesize diverse negation queries, and introduce a dynamic visual focus module for verification. Our evaluation reveals that advanced RS MLLMs struggle with negation, exhibiting hallucinations and substantial performance degradation. To close this gap, we propose NeFo, a novel test-time learning method that explicitly incorporates the logical role of negation into the model optimization. Remarkably, using about 5\% unlabeled test samples, NeFo significantly improves the negation understanding of models and shows strong generalization to unseen tasks.

遥感否定理解多模态测试时学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。