首个针对西班牙语文化偏见的系统评估,揭示大模型在拉美语境下的刻板印象问题。
SESGO: Spanish Evaluation of Stereotypical Generative Outputs
- 基于拉美文化特有表达设计新评测框架,覆盖性别、种族等四类社会偏见。
- 用4000+提示测试发现主流模型在西班牙语中存在显著偏见,且温度变化不影响偏见模式。
- 证明英语优化的去偏技术不适用于西班牙语,框架可扩展至其他语言文化。
本文针对多语言大模型中的偏见评估空白,聚焦拉丁美洲文化语境下的西班牙语。尽管全球广泛应用,现有评估仍以美国英语为中心,忽视其他语言与文化背景的潜在风险。我们提出一种新的、基于文化的偏见检测框架,通过改编BBQ数据集中的模糊提问法,融入四个社会维度(性别、种族、阶级、国籍)的区域性俗语和表达。利用超过4000个提示,构建结合准确率与错误方向的新指标,有效平衡模型性能与偏见对齐,在模糊与明确语境下均表现良好。据我们所知,这是首个系统性评估领先商用大模型对西班牙语文化偏见响应的研究,揭示了不同模型间偏见模式的差异。研究还表明,针对英语优化的去偏技术无法有效迁移至西班牙语任务,且偏见模式在不同采样温度下保持稳定。本模块化框架可自然拓展至新刻板印象、偏见类别或语言文化,为多元化语境中更公平、文化敏感的AI评估迈出关键一步。
原文摘要 · Abstract (English)
This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin American contexts. Despite widespread global deployment, current evaluations remain predominantly US-English-centric, leaving potential harms in other linguistic and cultural contexts largely underexamined. We introduce a novel, culturally-grounded framework for detecting social biases in instruction-tuned LLMs. Our approach adapts the underspecified question methodology from the BBQ dataset by incorporating culturally-specific expressions and sayings that encode regional stereotypes across four social categories: gender, race, socioeconomic class, and national origin. Using more than 4,000 prompts, we propose a new metric that combines accuracy with the direction of error to effectively balance model performance and bias alignment in both ambiguous and disambiguated contexts. To our knowledge, our work presents the first systematic evaluation examining how leading commercial LLMs respond to culturally specific bias in the Spanish language, revealing varying patterns of bias manifestation across state-of-the-art models. We also contribute evidence that bias mitigation techniques optimized for English do not effectively transfer to Spanish tasks, and that bias patterns remain largely consistent across different sampling temperatures. Our modular framework offers a natural extension to new stereotypes, bias categories, or languages and cultural contexts, representing a significant step toward more equitable and culturally-aware evaluation of AI systems in the diverse linguistic environments where they operate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。