多语言安全数据集存在显著语言差异,非洲语言覆盖严重不足。
Language-Specific Gaps in AI Safety Training Datasets
- 按语言切片审计21个资源,发现低资源语言数据缺口普遍存在
- 豪萨语数据质量低于自身翻译标准,而斯瓦希里语达标,证明问题可改进
- 自残与色情内容在非洲语言中无原生数据,非资源量能解释的结构性缺失
大型语言模型厂商常以覆盖十余种语言的多语言安全基准为由,声称其模型对非英语用户安全。我们发现这些整体覆盖声明在具体语言层面常不成立。对25个语言切片(涵盖豪萨语、斯瓦希里语、法语三类资源水平)中的21项资源进行审计,其中20项符合数据集定义。结果表明,来源透明度、标注可靠性、访问权限、危害分类覆盖和数据复用等问题在不同语言中反复出现,其模式部分但不完全对应资源层级。通过同管道对照实验,发现豪萨语切片输出未达自身论文设定的翻译质量阈值,而斯瓦希里语则轻松达标,证明这些差距可测量且可修复。此外,在所研究的非洲语言两级中,自残与性内容类别均无原生语言数据,这一完全缺失无法用资源量梯度解释。我们将其与已知的多语言越狱鲁棒性不对称现象(单轮攻击被抑制,多轮仍有效)联系起来,认为该现象与训练评估数据最薄弱的语言区域结构一致。本文贡献了可复用的切片级审计方法、跨层级实证比较及面向数据集创建者、模型厂商和会议的可验证多语言覆盖建议。数据集:https://huggingface.co/datasets/ChialukaOnuoha/safety-slice-audit
原文摘要 · Abstract (English)
Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users. We show that these collection-level coverage claims frequently do not survive inspection at the level of an individual language. Auditing 21 resources across 25 language slices, of which 20 count as datasets under our counting rules, spanning three languages chosen to represent low- (Hausa), mid- (Swahili), and high-resource (French) tiers, we find that gaps in provenance, annotation reliability, access, harm-taxonomy coverage, and data reuse recur in patterns that partially, but not fully, track resource level. Using a controlled within-pipeline comparison, we show a Hausa-language slice falling below its own paper's translation-quality acceptance threshold while the same pipeline's Swahili output clears the same bar comfortably; this is evidence that these gaps are measurable and addressable, not inherent. We further show that self-harm and sexual-content categories have no native-language coverage in either African-language tier we studied, a total rather than gradated gap that a purely resource-level account does not predict. We connect these findings to a documented, persistent asymmetry in multilingual jailbreak robustness (single-turn attacks largely mitigated, multi-turn attacks still effective), arguing that this asymmetry is structurally consistent with where our audit finds training and evaluation data thinnest. We contribute a reusable slice-level audit methodology, a cross-tier empirical comparison, and concrete recommendations for dataset creators, model providers, and venues aiming to make ``multilingual coverage'' claims verifiable rather than merely stated. Dataset: https://huggingface.co/datasets/ChialukaOnuoha/safety-slice-audit
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。