针对印尼多语言社会设计了评估大模型偏见的双轨基准
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages

- 构建双轨评测框架:深度对比对与广度生成结合
- 本地语言在宗教意识形态类偏见更严重,印尼语预训练数据源影响显著
- 适合研究跨文化公平性、本地化语言模型的学者和开发者
尽管印度尼西亚拥有1300多个族群和700多种本土语言,但大语言模型中的偏见研究仍不充分,导致在其独特多元的文化语言背景下,代表性公平性和地方刻板印象难以评估。为此,我们提出IndoBias,一个以文化为基础的偏见评测基准,用于评估印尼语及爪哇语、巽他语、望加锡语三种地方语言中大模型的偏见。IndoBias包含双轨评测:深度导向(对比对)与广度导向(基于社会学框架SPI、O*NET、WGI的生成)。结果显示,现有大模型(尤其是解码器模型)在印尼语中对典型句式存在强偏见;本地语言在意识形态与宗教类别中偏见更高。模型对不同地方实体的回应呈现非均匀刻板印象极性。此外,在印尼语中,通用网页文本(Common Crawl)比人工审校文本(如维基百科、新闻)引入更多偏见;而预训练中加入本地语言通常会加剧偏见。本工作强调在文化特定语境下研究偏见的重要性。警告:本文包含可能具有冒犯性、有害或偏见的数据示例。
原文摘要 · Abstract (English)
Despite being home to more than 1300 ethnic groups and 700 indigenous languages, bias in Large Language Models has not been fully studied in Indonesia, thus leaving a critical gap in evaluating representational fairness and localized stereotypes within its uniquely vast, multilingual, and diverse sociocultural landscape. To address this, we introduce IndoBias as a culturally-grounded bias benchmark to assess LLMs bias in Indonesian and three local languages: Javanese, Sundanese, and Makasar. IndoBias features dual perspective evaluation tracks: depth-oriented (with contrastive-pairs) and breadth-oriented (with generation-based), where the latter is grounded in social science frameworks (SPI, O*NET, and WGI). Our results show that existing LLMs -- particularly decoder models -- exhibit strong bias towards prototypical sentences in Indonesian, while local languages suffer higher bias under Ideology and Religion category. We also find that LLMs responses exhibit a non-uniform Stereotype Polarity when prompted with various local entities. Finally, we discover that, in Indonesian, Common Crawl texts introduce more bias during pretraining, compared to human-reviewed article texts (e.g., Wikipedia, News), whereas introducing local languages to pretraining generally increases bias. This work highlights the importance of studying bias in culture-specific context. Warning: This paper contains example data that may be offensive, harmful, or biased.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。