用人类与大模型协作构建西语文化偏见数据集,揭示多国差异。
Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration

- 人类与大模型协作生成候选偏见,本地化专家验证。
- 发现西语国家间存在显著文化差异的刻板印象表现。
- 框架可推广至其他语言,推动跨文化评测标准化。
大型语言模型中的偏见研究长期集中于英语语境,主要受限于非英语语料库的缺乏和小众文化手动标注成本过高。本文提出一种低成本的人类-大模型协同标注框架,并应用于构建涵盖欧洲与拉丁美洲多个西班牙语国家的西语刻板印象数据集EspanStereo。该数据集不仅包含已有文献记录的普遍性刻板印象,还捕捉到英语主导资源中缺失的文化特异性偏见。通过大模型生成候选刻板印象,由在地标注者进行验证,证明该框架能有效识别细微的区域性偏见。基于EspanStereo对西班牙语支持大模型的评估显示,不同国家间的刻板行为表现存在显著差异,凸显了更深入的文化根基评估的重要性。该框架可扩展至其他语言与地区,为多语言刻板印象基准提供可扩展路径。本工作拓展了大模型偏见分析的范围,奠定了全面跨文化偏见评估的基础。
原文摘要 · Abstract (English)
Research on stereotypes in large language models (LLMs) has largely focused on English-speaking contexts, due to the lack of datasets in other languages and the high cost of manual annotation in underrepresented cultures. To address this gap, we introduce a cost-efficient human-LLM collaborative annotation framework and apply it to construct EspanStereo, a Spanish-language stereotype dataset spanning multiple Spanish-speaking countries across Europe and Latin America. EspanStereo captures both well-documented stereotypes from prior literature and culturally specific biases absent from English-centric resources. Using LLMs to generate candidate stereotypes and in-culture annotators to validate them, we demonstrate the framework's effectiveness in identifying nuanced, region-specific biases. Our evaluation of Spanish-supporting LLMs using EspanStereo reveals significant variation in stereotypical behavior across countries, highlighting the need for more culturally grounded assessments. Beyond Spanish, our framework is adaptable to other languages and regions, offering a scalable path toward multilingual stereotype benchmarks. This work broadens the scope of stereotype analysis in LLMs and lays the groundwork for comprehensive cross-cultural bias evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。