arXiv:2510.26861cs.IRcs.CL2025-10ACL

发现多模态检索中语言与文化偏见导致结果不公,提出新评测基准。

Evaluating Perspectival Biases in Cross-Modal Retrieval

  • 构建3XCM基准,分离语言与文化对检索的影响。
  • 低资源语言下,视觉相似性被文化熟悉度主导。
  • 需显式解耦语言与文化,才能实现公平检索。

多模态检索系统本应基于语义空间,不受查询语言或文化背景影响。但实践中,检索结果系统性地反映视角偏见:由语言流行度和文化关联性塑造的偏差。我们提出跨文化、跨模态、跨语言多模态(3XCM)评测基准,以分离这些影响。研究显示,在图像到文本检索中,模型更倾向流行语言条目而非语义准确匹配;在文本到图像检索中,联合嵌入空间存在“拉拽效应”,即语义对齐与语言相关的文化关联相互干扰。当语义表示不足时,尤其在低资源语言中,相似性逐渐由文化熟悉的视觉模式主导,造成系统性关联偏见。结果表明,实现公平多模态检索需针对性策略,显式解耦语言与文化,而非仅依赖数据规模扩大。该工作强调,语言与文化偏见应被视为多模态表征学习中可测量的独立挑战。

原文摘要 · Abstract (English)

Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practice, however, retrieval outcomes systematically reflect perspectival biases: deviations shaped by linguistic prevalence and cultural associations. We introduce the Cross-Cultural, Cross-Modal, Cross-lingual Multimodal (3XCM) benchmark to isolate these effects. Results from our studies indicate that, for image-to-text retrieval, models tend to favor entries from prevalent languages over those that are semantically faithful. For text-to-image retrieval, we observe a consistent "tugging effect" in the joint embedding space between semantic alignment and language-conditioned cultural association. When semantic representations are insufficiently resolved, particularly in low-resource languages, similarity is increasingly governed by culturally familiar visual patterns, leading to systematic association bias in retrieval. Our findings suggest that achieving equitable multimodal retrieval necessitates targeted strategies that explicitly decouple language from culture, rather than relying solely on broader data exposure. This work highlights the need to treat linguistic and cultural biases as distinct, measurable challenges in multimodal representation learning.

多模态检索文化偏见语言公平嵌入空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。