arXiv:2608.00877cs.LG2026-08

用可验证性筛选地理信息,让遥感多模态大模型更可信。

GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs

论文配图:GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs
图 1 · 摘自论文原文
  • 只注入图像无法验证的地理事实,避免与视觉证据冲突。
  • 在fMoW数据集上提升土地利用识别准确率12.06至17.19点。
  • 不依赖训练,适合希望提升模型可信度的研究者。

遥感多模态大语言模型常做出图像无法证实的事实判断,如设施身份或功能。通过坐标索引地理检索可补充此类知识,使三个开源模型在fMoW土地利用任务上的准确率提升12.06至17.19个百分点。然而,检索记录可能与可见证据矛盾,我们发现模型常盲目采纳记录而非遵循图像。因此,我们提出基于跨模态可验证性的信任机制:地理记录对图像无法验证的属性最有用,对可验证属性则最危险。为此设计GeoArbiter,一种无需训练的流水线,仅注入图像不可验证的地理事实。相比仲裁提示,内容级过滤保留84.69至87.15%的完整检索增益,使盲源评判下的声明级幻觉降低9.58至26.34%,且在三种模型上均增强对冲突记录的鲁棒性。结果表明,基于可验证性的内容选择是融合不可靠地理知识的有效方法。

原文摘要 · Abstract (English)

Remote-sensing multimodal large language models (MLLMs) often assert facts that imagery cannot establish, such as a facility's identity or function. Coordinate-keyed geographic retrieval can supply this missing knowledge, improving fMoW land-use accuracy by 12.06--17.19 points across three open MLLMs. However, retrieved records can also contradict visible evidence, and we find that models frequently follow the records even when the image is decisive. We argue that source trust should therefore depend on \emph{cross-modal verifiability}: geographic records are most useful for attributes the image cannot verify and most dangerous when they dispute visually verifiable attributes. We introduce GeoArbiter, a training-free pipeline that operationalizes this principle by injecting only image-unverifiable geographic facts. Unlike arbitration prompts, which leak across attribute types and bias yes/no responses, content-level filtering preserves 84.69--87.15\% of the full-retrieval accuracy gain, reduces claim-level hallucination by 9.58--26.34\% under a source-blinded judge, and improves robustness to conflicting records across all three models. These results identify verifiability-guided content selection as a simple, effective mechanism for grounding remote-sensing MLLMs in fallible geographic knowledge.

遥感多模态可信生成知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。