测试AI对25国文化禁忌手势的识别能力,发现多模型存在偏见。
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
- 构建288组手势-国家数据集,标注文化敏感性与语境。
- 美式偏见明显:文本生成图像模型在非美国场景误判率超40%。
- 大模型易过度报警,视觉语言模型常推荐不恰当手势。
手势是非语言交流的重要组成部分,其含义因文化而异,误解可能带来严重的社会与外交后果。随着AI系统在全球应用中的深入集成,确保其不无意中传播文化冒犯至关重要。为此,我们提出多文化不当手势与非语言符号数据集(MC-SIGNS),包含25种手势与85个国家的288个手势-国家组合,涵盖冒犯性、文化意义及语境因素的标注。通过系统评估,我们发现:文本到图像(T2I)系统表现出显著的美国中心偏见,在美国语境下识别准确率高于非美国语境;大型语言模型(LLMs)倾向于过度标记手势为冒犯;视觉-语言模型(VLMs)在回应如‘祝你好运’等通用概念时,常默认采用美国式解读,频繁建议文化不合适的动作。这些结果凸显了建立文化敏感型AI安全机制的紧迫性,以保障AI技术在全球范围内的公平部署。
原文摘要 · Abstract (English)
Gestures are an integral part of non-verbal communication, with meanings that vary across cultures, and misinterpretations that can have serious social and diplomatic consequences. As AI systems become more integrated into global applications, ensuring they do not inadvertently perpetuate cultural offenses is critical. To this end, we introduce Multi-Cultural Set of Inappropriate Gestures and Nonverbal Signs (MC-SIGNS), a dataset of 288 gesture-country pairs annotated for offensiveness, cultural significance, and contextual factors across 25 gestures and 85 countries. Through systematic evaluation using MC-SIGNS, we uncover critical limitations: text-to-image (T2I) systems exhibit strong US-centric biases, performing better at detecting offensive gestures in US contexts than in non-US ones; large language models (LLMs) tend to over-flag gestures as offensive; and vision-language models (VLMs) default to US-based interpretations when responding to universal concepts like wishing someone luck, frequently suggesting culturally inappropriate gestures. These findings highlight the urgent need for culturally-aware AI safety mechanisms to ensure equitable global deployment of AI technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。