arXiv:2602.05437cs.CL2026-02ACL被引 1

多语言视觉模型虽答对问题,却常接受错误的反事实解释,尤其在阿拉伯语方言中。

Once Correct, Still Wrong: Counterfactual Hallucination in Multilingual Vision-Language Models

  • 构建跨文化多模态数据集M²CQA,覆盖17个中东北非国家图像与多语言陈述
  • 发现阿拉伯语方言中反事实幻觉率(CFHR)显著上升,即使正确率仍高
  • 先推理会加剧幻觉,先回答再解释能提升模型鲁棒性

视觉语言模型(VLMs)在保持高准确率的同时,仍可能接受文化上合理但视觉上错误的解释。现有幻觉评估基准很少测试这一失败模式,尤其缺乏非西方语境和非英语数据。我们提出M²CQA,一个基于17个中东北非(MENA)国家图像的多模态文化基准,包含英文、阿拉伯语及其方言的对比真实与反事实陈述。为分离幻觉与准确率,我们引入反事实幻觉率(CFHR),衡量在正确回答真实陈述前提下接受反事实的程度。在多种提示策略下评估先进VLMs,发现阿拉伯语中(尤其是方言)的CFHR急剧上升,即使真实陈述准确率维持高位。此外,先推理的提示方式会持续增加反事实幻觉,而先回答后解释则提升鲁棒性。数据集已公开供社区使用(https://huggingface.co/datasets/QCRI/M2CQA)。

原文摘要 · Abstract (English)

Vision-language models (VLMs) can achieve high accuracy while still accepting culturally plausible but visually incorrect interpretations. Existing hallucination benchmarks rarely test this failure mode, particularly outside Western contexts and English. We introduce M$^2$CQA, a culturally grounded multimodal benchmark built from images spanning 17 MENA countries, paired with contrastive true and counterfactual statements in English, Arabic, and its dialects. To isolate hallucination beyond raw accuracy, we propose the CounterFactual Hallucination Rate (CFHR), which measures counterfactual acceptance conditioned on correctly answering the true statement. Evaluating state-of-the-art VLMs under multiple prompting strategies, we find that CFHR rises sharply in Arabic, especially in dialects, even when true-statement accuracy remains high. Moreover, reasoning-first prompting consistently increases counterfactual hallucination, while answering before justifying improves robustness. We make the dataset publicly available for the community (https://huggingface.co/datasets/QCRI/M2CQA)).

多模态幻觉检测阿拉伯语文化偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。