arXiv:2602.18729cs.CVcs.AI2026-02Conference of the …

构建安全与文化细粒度对齐基准,揭示视觉语言模型在微小差异下的理解缺陷

MiSCHiEF: A Benchmark in Minimal-Pairs of Safety and Culture for Holistic Evaluation of Fine-Grained Image-Caption Alignment

  • 设计最小差异图像-标题对,对比安全与文化场景中的细微语义差别
  • 模型更擅长确认正确配对,拒错能力明显不足,跨模态对齐仍存挑战
  • 适合评估模型在社会敏感场景下的细粒度理解能力,如风险识别与文化差异判断

细粒度图像-标题对齐对视觉语言模型(VLMs)至关重要,尤其在识别真实世界风险或区分文化符号等社会敏感场景中。本文提出 MiSCHiEF 基准,包含基于最小差异对设计的两个数据集:安全领域(MiS)与文化领域(MiC)。每个样本包含两组极相似的图像和标题,分别表示安全/不安全场景或不同文化背景下的文化代理。在两项任务中评估四种 VLMs,发现模型在确认正确配对时表现优于拒绝错误配对;且从高度相似标题中选出正确标题比反向任务表现更好。整体结果表明当前 VLMs 在跨模态精确定位上仍存在持续性对齐难题,凸显了在细微语义与视觉差异下精准对齐的难度。

原文摘要 · Abstract (English)

Fine-grained image-caption alignment is crucial for vision-language models (VLMs), especially in socially critical contexts such as identifying real-world risk scenarios or distinguishing cultural proxies, where correct interpretation hinges on subtle visual or linguistic clues and where minor misinterpretations can lead to significant real-world consequences. We present MiSCHiEF, a set of two benchmarking datasets based on a contrastive pair design in the domains of safety (MiS) and culture (MiC), and evaluate four VLMs on tasks requiring fine-grained differentiation of paired images and captions. In both datasets, each sample contains two minimally differing captions and corresponding minimally differing images. In MiS, the image-caption pairs depict a safe and an unsafe scenario, while in MiC, they depict cultural proxies in two distinct cultural contexts. We find that models generally perform better at confirming the correct image-caption pair than rejecting incorrect ones. Additionally, models achieve higher accuracy when selecting the correct caption from two highly similar captions for a given image, compared to the converse task. The results, overall, highlight persistent modality misalignment challenges in current VLMs, underscoring the difficulty of precise cross-modal grounding required for applications with subtle semantic and visual distinctions.

视觉语言模型细粒度对齐安全评估文化差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。