提出跨模态一致性测试,评估视觉语言模型在图文匹配中的内在自洽性。
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
- 通过文本、图像或双模态提问,检验模型对场景相似性的判断一致性
- 不依赖真实答案,只关注模型输出是否自洽,而非绝对正确
- 适合研究模型幻觉与多模态对齐问题的研究者使用
视觉语言模型(VLMs)通常通过场景感知的视觉问答(VQA)任务进行评估,认为良好的VQA表现可预示其在更广泛多模态任务中的潜力。然而,场景感知的VQA未能充分捕捉输入偏差或由模态间错位引发的幻觉问题。为此,本文提出跨模态对齐相似性测试(CAST),用于探测VLM在不同模态间的一致性。该测试要求模型仅通过文本、仅通过图像或两者结合来识别两个场景间的相似性,并评估其生成的相似性描述的真实性。由于无标准答案可供比对,评估重点并非客观准确性,而是模型输出是否内部一致。我们主张:尽管并非所有自洽的模型都具备能力或准确,但所有具备能力的VLM必须具备自洽性。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) are typically evaluated with Visual Question Answering (VQA) tasks which assess a model's understanding of scenes. Good VQA performance is taken as evidence that the model will perform well on a broader range of tasks that require both visual and language inputs. However, scene-aware VQA does not fully capture input biases or assess hallucinations caused by a misalignment between modalities. To address this, we propose a Cross-modal Alignment Similarity Test (CAST) to probe VLMs for self-consistency across modalities. This test involves asking the models to identify similarities between two scenes through text-only, image-only, or both and then assess the truthfulness of the similarities they generate. Since there is no ground-truth to compare against, this evaluation does not focus on objective accuracy but rather on whether VLMs are internally consistent in their outputs. We argue that while not all self-consistent models are capable or accurate, all capable VLMs must be self-consistent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。