构建首个印尼语文化安全评估数据集,提升本地化大模型安全性
IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
- 基于印尼社会文化设计安全分类体系,覆盖五种语言变体
- 现有印尼大模型在口语和地方语言中易生成不安全内容
- 微调后显著提升安全性,适合多语言场景负责任部署
尽管区域性大语言模型日益发展,其安全性在文化多元地区仍研究不足,尤其在印尼这类重视本地规范的社区中尤为关键。本文提出IndoSafety,首个面向印尼语环境的高质量、人工验证的安全评估数据集,涵盖正式与口语化的印尼语,以及爪哇语、巽他语和明古鲁语三种主要地方语言。通过扩展既有安全框架,构建契合印尼社会文化语境的分类体系。实验发现,现有印尼语大模型在口语及地方语言场景中常生成不安全内容;而使用IndoSafety进行微调后,安全性能显著提升,同时保持任务表现。本工作凸显了文化适配性安全评估的重要性,并为多语言环境下负责任的大模型部署提供具体路径。警告:本文包含可能具有冒犯性、危害性或偏见的示例数据。
原文摘要 · Abstract (English)
Although region-specific large language models (LLMs) are increasingly developed, their safety remains underexplored, particularly in culturally diverse settings like Indonesia, where sensitivity to local norms is essential and highly valued by the community. In this work, we present IndoSafety, the first high-quality, human-verified safety evaluation dataset tailored for the Indonesian context, covering five language varieties: formal and colloquial Indonesian, along with three major local languages: Javanese, Sundanese, and Minangkabau. IndoSafety is constructed by extending prior safety frameworks to develop a taxonomy that captures Indonesia's sociocultural context. We find that existing Indonesian-centric LLMs often generate unsafe outputs, particularly in colloquial and local language settings, while fine-tuning on IndoSafety significantly improves safety while preserving task performance. Our work highlights the critical need for culturally grounded safety evaluation and provides a concrete step toward responsible LLM deployment in multilingual settings. Warning: This paper contains example data that may be offensive, harmful, or biased.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。