arXiv:2511.04090cs.SIcs.AI2025-11

针对拉美文化表达不足,构建本地化数据集提升大模型文化敏感性。

Advancing Equitable AI: Evaluating Cultural Expressiveness in LLMs for Latin American Contexts

  • 基于拉美历史与社会背景构建新数据集,突破西方中心范式。
  • 部分模型文化表达能力显著更优,但多数存在情感偏差(p<0.001)。
  • 微调后Mistral-7B文化表达提升42.9%,适合关注公平AI的研究者。

当前人工智能系统常反映经济发达地区偏见,因数据分布不均而边缘化拉丁美洲等发展中地区。本文考察了大模型对多元拉美语境的表征,揭示了发达与发展中地区数据间的差异。英语主导地位强化了对西班牙语、葡萄牙语及克丘亚语、纳瓦特尔语等原住民语言的压制,使拉美视角被西方框架所塑造。为此,我们提出一个融合拉美历史与社会政治背景的文化敏感数据集,挑战欧洲中心主义模型。在测试文化情境感知能力时,采用新型文化表达度量、统计检验与语言分析评估六种语言模型。结果显示,某些模型更能捕捉拉美视角,另一些则存在显著情感错位(p < 0.001)。用该数据集微调Mistral-7B,其文化表达能力提升42.9%。我们倡导以反映拉美历史、原住民知识与多语言生态的数据集推动公平AI发展,并强调以社区为中心的方法来增强被边缘化声音的表达。

原文摘要 · Abstract (English)

Artificial intelligence (AI) systems often reflect biases from economically advanced regions, marginalizing contexts in economically developing regions like Latin America due to imbalanced datasets. This paper examines AI representations of diverse Latin American contexts, revealing disparities between data from economically advanced and developing regions. We highlight how the dominance of English over Spanish, Portuguese, and indigenous languages such as Quechua and Nahuatl perpetuates biases, framing Latin American perspectives through a Western lens. To address this, we introduce a culturally aware dataset rooted in Latin American history and socio-political contexts, challenging Eurocentric models. We evaluate six language models on questions testing cultural context awareness, using a novel Cultural Expressiveness metric, statistical tests, and linguistic analyses. Our findings show that some models better capture Latin American perspectives, while others exhibit significant sentiment misalignment (p < 0.001). Fine-tuning Mistral-7B with our dataset improves its cultural expressiveness by 42.9%, advancing equitable AI development. We advocate for equitable AI by prioritizing datasets that reflect Latin American history, indigenous knowledge, and diverse languages, while emphasizing community-centered approaches to amplify marginalized voices.

公平AI文化表达大模型拉美

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。