针对日语大模型的文化偏见,设计了本土化对抗测试集。
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
- 用日语母语者手工构建对话场景,暴露模型隐性偏见。
- 九款日语模型平均准确率仅23%,远低于随机水平。
- 适合研究日语AI偏见、文化适配与伦理评估的学者。
语言中的社会偏见根植于文化规范,不同地区差异显著,导致刻板印象表现多样。现有非英语语境下大语言模型(LLMs)的社会偏见评估多依赖英文基准的翻译,未能反映本地文化,如日本特有的等级关系、方言差异和传统性别角色。为此,我们提出日语文化对抗偏见基准JUBAKU,通过人工精心设计的对抗构造,揭示十类文化范畴中的潜在偏见。与现有基准不同,JUBAKU由日语母语标注者手写对话场景,专门用于触发并暴露日本LLMs的隐性社会偏见。我们在九款日本LLMs及三款基于英文基准适配的模型上进行评估,所有模型在JUBAKU上的表现均低于50%随机基线,平均准确率为23%(范围13%至33%),尽管在其他基准上表现更高。人类标注者识别无偏响应的准确率达91%,验证了JUBAKU的可靠性及其对模型的对抗性。
原文摘要 · Abstract (English)
Social biases reflected in language are inherently shaped by cultural norms, which vary significantly across regions and lead to diverse manifestations of stereotypes. Existing evaluations of social bias in large language models (LLMs) for non-English contexts, however, often rely on translations of English benchmarks. Such benchmarks fail to reflect local cultural norms, including those found in Japanese. For instance, Western benchmarks may overlook Japan-specific stereotypes related to hierarchical relationships, regional dialects, or traditional gender roles. To address this limitation, we introduce Japanese cUlture adversarial BiAs benchmarK Under handcrafted creation (JUBAKU), a benchmark tailored to Japanese cultural contexts. JUBAKU uses adversarial construction to expose latent biases across ten distinct cultural categories. Unlike existing benchmarks, JUBAKU features dialogue scenarios hand-crafted by native Japanese annotators, specifically designed to trigger and reveal latent social biases in Japanese LLMs. We evaluated nine Japanese LLMs on JUBAKU and three others adapted from English benchmarks. All models clearly exhibited biases on JUBAKU, performing below the random baseline of 50% with an average accuracy of 23% (ranging from 13% to 33%), despite higher accuracy on the other benchmarks. Human annotators achieved 91% accuracy in identifying unbiased responses, confirming JUBAKU's reliability and its adversarial nature to LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。