构建跨模态语义等价评测基准,揭示视觉语言模型在多模态推理中的系统性偏差。
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- 设计四领域语义等价数据对,用不同符号系统避免图像文字匹配干扰
- 21个模型测试显示视觉性能普遍低于语言,跨模态一致率低
- 发现文本分词错误与视觉幻觉是主要错误来源,结果对图像变换稳健
评估视觉语言模型(VLMs)在不同表示间是否保持一致推理能力具有挑战性,因模态对比常被任务差异和信息不对称混淆。我们提出SEAM,一个涵盖四个具备标准化文本与视觉表示领域的语义等价跨模态基准。通过在不同模态使用异构符号系统,而非基于OCR的图像-文本配对,SEAM严格评估VLMs在文本-符号与视觉-空间推理方面的能力。在21个当代模型上,我们观察到系统性模态失衡:尽管问题包含语义等价信息,视觉性能普遍落后于语言,且跨模态一致性较低。误差分析揭示两大主因:领域符号表示下的文本感知失败(源于分词),以及引发幻觉的视觉感知失败。此外,结果在多种视觉变换下仍保持稳健。SEAM为测量与提升无模态偏见推理提供了受控语境。
原文摘要 · Abstract (English)
Evaluating whether vision-language models (VLMs) reason consistently across representations is challenging because modality comparisons are typically confounded by task differences and asymmetric information. We introduce SEAM, a benchmark that pairs semantically equivalent inputs across four domains that have existing standardized textual and visual notations. By employing distinct notation systems across modalities, in contrast to OCR-based image-text pairing, SEAM provides a rigorous comparative assessment of the textual-symbolic and visual-spatial reasoning capabilities of VLMs. Across 21 contemporary models, we observe systematic modality imbalance: vision frequently lags language in overall performance, despite the problems containing semantically equivalent information, and cross-modal agreement is relatively low. Our error analysis reveals two main drivers: textual perception failures from tokenization in domain notation and visual perception failures that induce hallucinations. We also show that our results are largely robust to visual transformations. SEAM establishes a controlled, semantically equivalent setting for measuring and improving modality-agnostic reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。