arXiv:2608.04244cs.CVcs.CL2026-08

测试多模态模型在图文冲突时如何判断,发现多数模型易被错误文本误导。

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

论文配图:SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
图 1 · 摘自论文原文
  • 设计五类图像变体,精准控制图文冲突场景。
  • 图文冲突使定位误差平均升高4.8倍,最高达1347公里。
  • 即使基础定位准确,模型仍可能被错误文本引导偏离真实位置。

多模态大语言模型(MLLMs)通过融合视觉与文本线索进行真实场景的有根据预测,但现有基准很少揭示它们在证据冲突时如何权衡。我们提出SIGNPOST-Bench,一个受控的反事实基准,用于评估图文冲突解决能力。每张源图像生成原始、空白、相似、随机和对抗性五类变体。通过合成局部化场景-文本干预,保留非文本内容,实现对定位性能变化与文本引导目标偏移的成对测量。SIGNPOST-Bench包含5,111组反事实样本和25,555个图像变体,来自四个数据集。我们评估了七家厂商的20个MLLM。相比原始图像,对抗性变体使中位定位误差从282公里升至1,347公里,提升4.8倍。在可地理定位的对抗样本中,6.5%-20.1%的预测结果距离注入目标小于50公里,且所有模型在空白到对抗图像间均表现出目标距离的平均正向减少。一致、无关和冲突文本替换产生不同影响,而干净输入下的定位性能无法完全预测模型对冲突文本的鲁棒性。这些结果将视觉地理定位作为连续诊断工具,为评估MLLM解决多模态矛盾证据提供了可控框架。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.

多模态模型图文冲突基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。