轻量级多语言内容审核模型,适配新加坡本地语种
LionGuard 2: Building Lightweight, Data-Efficient & Localised Multilingual Content Moderators
- 基于预训练嵌入与序数分类器,无需微调大模型
- 在17个基准上超越多个商用开源系统
- 适合政府级部署,支持英/中/马/部分泰米尔语
现代内容审核系统日益支持多语言,但常忽略本地化和低资源变体,导致实际部署存在安全漏洞。小型模型虽可替代大语言模型,但仍需大量数据和算力。我们提出LionGuard 2,一种面向新加坡场景的轻量级、多语言内容审核分类器,支持英语、中文、马来语及部分泰米尔语。该模型基于预训练OpenAI嵌入与多头序数分类器,在17个基准测试中表现优于多个商业及开源系统,涵盖新加坡本地与公开英文数据集。系统已在新加坡政府内部实际部署,验证了其大规模应用的有效性。研究发现,高质量本地数据与稳健的多语言嵌入可在不微调大模型的前提下实现优异审核性能。我们已开放模型权重及部分训练数据,以推动大模型安全研究。
原文摘要 · Abstract (English)
Modern moderation systems increasingly support multiple languages, but often fail to address localisation and low-resource variants - creating safety gaps in real-world deployments. Small models offer a potential alternative to large LLMs, yet still demand considerable data and compute. We present LionGuard 2, a lightweight, multilingual moderation classifier tailored to the Singapore context, supporting English, Chinese, Malay, and partial Tamil. Built on pre-trained OpenAI embeddings and a multi-head ordinal classifier, LionGuard 2 outperforms several commercial and open-source systems across 17 benchmarks, including both Singapore-specific and public English datasets. The system is actively deployed within the Singapore Government, demonstrating practical efficacy at scale. Our findings show that high-quality local data and robust multilingual embeddings can achieve strong moderation performance, without fine-tuning large models. We release our model weights and part of our training data to support future work on LLM safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。