构建非洲四国的多语言刻板印象数据集,助力AI安全评估
SAFARI: A Community-Engaged Approach and Dataset of Stereotype Resources in the Sub-Saharan African Context
- 通过电话调查等社区参与方式收集数据
- 涵盖3534条英文与3206条本土语言刻板印象
- 适合关注AI公平性与非洲语料研究者
刻板印象资源对评估生成式AI模型的安全性至关重要,但当前全球覆盖不足。本文针对撒哈拉以南非洲地区严重缺乏NLP资源的问题,聚焦加纳、肯尼亚、尼日利亚和南非四个国家,采用社会文化情境化、社区参与的方法,包括使用母语进行电话调查,建立一种可复现且敏感于该地区复杂语言多样性和传统口头文化的采集方法。通过在不同族裔和人口背景间均衡采样,确保广泛覆盖,最终建成包含3,534条英文刻板印象和3,206条15种本土语言刻板印象的数据集。
原文摘要 · Abstract (English)
Stereotype repositories are critical to assess generative AI model safety, but currently lack adequate global coverage. It is imperative to prioritize targeted expansion, strategically addressing existing deficits, over merely increasing data volume. This work introduces a multilingual stereotype resource covering four sub-Saharan African countries that are severely underrepresented in NLP resources: Ghana, Kenya, Nigeria, and South Africa. By utilizing socioculturally-situated, community-engaged methods, including telephonic surveys moderated in native languages, we establish a reproducible methodology that is sensitive to the region's complex linguistic diversity and traditional orality. By deliberately balancing the sample across diverse ethnic and demographic backgrounds, we ensure broad coverage, resulting in a dataset of 3,534 stereotypes in English and 3,206 stereotypes across 15 native languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。