测试日语大模型在刻板印象提示下的安全表现,发现其拒答率低且易生成有害内容。
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
- 构建3612个日语刻板印象提示,测试三类语言模型响应
- 日语原生模型拒答率最低,更易生成负面和有毒内容
- 提示格式显著影响输出,跨模型存在群体偏见差异
近年来,大语言模型虽展现巨大潜力,但其内在刻板印象与偏见引发的安全问题日益突出。现有研究多依赖间接评估方法,通过让模型在关联特定社会群体的句子对中进行选择。近期直接评估方法兴起,通过开放性响应克服先前方法的标注者偏见等局限。然而,多数研究聚焦英文模型,非英语模型尤其是日语模型的研究仍较少。本研究在直接评估框架下,分析日语大模型对刻板印象触发提示的安全性。我们结合301个按年龄、性别等属性分类的社会群体术语与12种刻板印象诱导模板,构建了3,612个日语提示。分析了三个基础模型(分别基于日语、英语、中文训练)的响应。结果表明,日语原生模型LLM-jp拒答率最低,更可能生成有毒和负面回应。此外,提示格式显著影响所有模型输出,生成内容对特定社会群体呈现夸张反应,且不同模型间存在差异。研究揭示日语大模型伦理安全机制不足,即使高精度模型在处理日语提示时亦会生成偏见输出。呼吁加强日语大模型的安全机制与偏见缓解策略,推动跨语言边界的人工智能伦理讨论。
原文摘要 · Abstract (English)
In recent years, Large Language Models have attracted growing interest for their significant potential, though concerns have rapidly emerged regarding unsafe behaviors stemming from inherent stereotypes and biases. Most research on stereotypes in LLMs has primarily relied on indirect evaluation setups, in which models are prompted to select between pairs of sentences associated with particular social groups. Recently, direct evaluation methods have emerged, examining open-ended model responses to overcome limitations of previous approaches, such as annotator biases. Most existing studies have focused on English-centric LLMs, whereas research on non-English models, particularly Japanese, remains sparse, despite the growing development and adoption of these models. This study examines the safety of Japanese LLMs when responding to stereotype-triggering prompts in direct setups. We constructed 3,612 prompts by combining 301 social group terms, categorized by age, gender, and other attributes, with 12 stereotype-inducing templates in Japanese. Responses were analyzed from three foundational models trained respectively on Japanese, English, and Chinese language. Our findings reveal that LLM-jp, a Japanese native model, exhibits the lowest refusal rate and is more likely to generate toxic and negative responses compared to other models. Additionally, prompt format significantly influence the output of all models, and the generated responses include exaggerated reactions toward specific social groups, varying across models. These findings underscore the insufficient ethical safety mechanisms in Japanese LLMs and demonstrate that even high-accuracy models can produce biased outputs when processing Japanese-language prompts. We advocate for improving safety mechanisms and bias mitigation strategies in Japanese LLMs, contributing to ongoing discussions on AI ethics beyond linguistic boundaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。