arXiv:2604.18942cs.CL2026-04

跨语言视觉模型在否定理解上存在偏见,新基准揭示公平性差异

Disparities In Negation Understanding Across Languages In Vision-Language Models

  • 构建首个多语言否定理解人工验证基准,覆盖7种语系
  • 非拉丁字母语言中标准CLIP性能接近随机,多语言模型表现更优
  • 纠正方案对不同语言效果不一,凸显语言结构影响公平性

视觉语言模型普遍存在肯定偏见:即使正确描述包含否定(如“无X”),仍倾向于选择肯定表述(如“X存在”)。尽管英语中已有相关研究与解决方案,但不同语言的否定表达受词形、语序和缩合词等影响各异,导致现有方法可能无法公平适用于所有语言社区。本文首次提出一个经人工验证的多语言否定理解基准,涵盖七种语系多样语言:英语、中文、阿拉伯语、希腊语、俄语、他加禄语和西班牙语。评估了三种视觉语言模型(CLIP、SigLIP、MultiCLIP),发现标准CLIP在非拉丁文字语言中表现接近随机,而MultiCLIP在各语言中均取得最高且最一致的准确率。同时测试了提出的空间修正模型SpaceVLM,结果显示其在英语、希腊语、西班牙语和他加禄语中显著提升性能,但在不同类型语言间效果差异明显。该结果表明,语言特征(如形态、书写系统、否定结构)与模型改进之间存在交互关系,影响公平性。随着视觉语言模型全球部署,建立多语言基准对理解解决方案的实际适用范围至关重要。

原文摘要 · Abstract (English)

Vision-language models (VLMs) exhibit affirmation bias: a systematic tendency to select positive captions ("X is present") even when the correct description contains negation ("no X"). While prior work has documented this failure mode in English and proposed solutions, negation manifests differently across languages through varying morphology, word order, and cliticization patterns, raising the question of whether these solutions serve all linguistic communities equitably. We introduce the first human-verified multilingual negation benchmark, spanning seven typologically diverse languages: English, Mandarin Chinese, Arabic, Greek, Russian, Tagalog, and Spanish. Evaluating three VLMs - CLIP, SigLIP, and MultiCLIP - we find that standard CLIP performs at or below chance on non-Latin-script languages, while MultiCLIP achieves the highest and most uniform accuracy. We also evaluate SpaceVLM, a proposed negation correction, and find that it produces substantial improvements for several languages - particularly English, Greek, Spanish, and Tagalog - while showing varied effectiveness across typologically different languages. This variation reveals that linguistic properties like morphology, script, and negation structure interact with model improvements in fairness-relevant ways. As VLMs are deployed globally, multilingual benchmarks are essential for understanding not just whether solutions work, but for whom.

视觉语言模型跨语言否定理解公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。