为东南亚打造本土化AI安全防护模型,兼顾文化差异与多语言适配。
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
- 构建代理式数据生成框架,自动生成符合当地文化的高质量安全数据。
- 在多语种评测中显著优于现有模型,对地区敏感内容检测准确率更高。
- 适合关注区域AI伦理、跨文化部署的开发者与政策制定者使用。
具备文化意识的安全防护机制在真实场景中至关重要,安全不仅关乎常识,更涉及多样化的本地价值观、行为规范和区域性法规。然而,受限于资源匮乏和母语标注者稀缺,构建大规模、文化契合的数据集极具挑战,导致多数防护模型依赖英文数据的机器翻译,难以捕捉地域与文化细节。本文提出一种新型代理式数据生成框架,可规模化生成真实、区域特定的安全数据,用于东南亚(SEA)场景。基于此,我们推出SEA-Guard系列,首个扎根于东南亚文化背景的多语言安全防护模型。在多个基准与文化变体测试中,SEA-Guard在识别地区敏感或有害内容方面持续领先,同时保持优异的通用安全性能。
原文摘要 · Abstract (English)
Culturally aware safeguards are crucial for AI alignment in real-world settings, where safety extends beyond common sense and encompasses diverse local values, norms, and region-specific regulations. However, building large-scale, culturally grounded datasets is challenging due to limited resources and a scarcity of native annotators. Consequently, many safeguard models rely on machine translation of English datasets, often missing regional and cultural nuances. We present a novel agentic data-generation framework to scalably create authentic, region-specific safety datasets for Southeast Asia (SEA). On this foundation, we introduce the SEA-Guard family, the first multilingual safeguard models grounded in SEA cultural contexts. Evaluated across multiple benchmarks and cultural variants, SEA-Guard consistently outperforms existing safeguards at detecting regionally sensitive or harmful content while maintaining strong general safety performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。