arXiv:2503.16826cs.CL2025-03被引 5

测试多模态模型在跨文化场景下的偏见,发现对低资源文化识别准确率下降超58%。

When Tom Eats Kimchi: Evaluating Cultural Bias of Multimodal Large Language Models in Cultural Mixture Contexts

  • 构建跨文化基准MixCuBe,评估模型在五国四族背景下的表现。
  • 低资源文化下模型准确率下降最多达58%,高资源文化则更稳定。
  • 揭示当前多模态大模型对非主流文化的系统性偏差,适合关注AI公平性的研究者参考。

在全球化背景下,多模态大语言模型(MLLMs)需正确识别混合文化输入。例如,无论亚洲女性还是非洲男性食用泡菜,模型都应准确识别为韩国食物。然而,现有模型过度依赖人物视觉特征,导致误判。为此,我们提出跨文化偏见基准MixCuBe,涵盖五国四族元素。结果显示,高资源文化下模型准确率更高且对扰动不敏感,而低资源文化中表现显著下降。最佳模型GPT-4o在低资源文化中原始与扰动设置间准确率差异高达58%。数据集已公开于:https://huggingface.co/datasets/kyawyethu/MixCuBe。

原文摘要 · Abstract (English)

In a highly globalized world, it is important for multi-modal large language models (MLLMs) to recognize and respond correctly to mixed-cultural inputs. For example, a model should correctly identify kimchi (Korean food) in an image both when an Asian woman is eating it, as well as an African man is eating it. However, current MLLMs show an over-reliance on the visual features of the person, leading to misclassification of the entities. To examine the robustness of MLLMs to different ethnicity, we introduce MixCuBe, a cross-cultural bias benchmark, and study elements from five countries and four ethnicities. Our findings reveal that MLLMs achieve both higher accuracy and lower sensitivity to such perturbation for high-resource cultures, but not for low-resource cultures. GPT-4o, the best-performing model overall, shows up to 58% difference in accuracy between the original and perturbed cultural settings in low-resource cultures. Our dataset is publicly available at: https://huggingface.co/datasets/kyawyethu/MixCuBe.

多模态模型文化偏见公平性评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。