arXiv:2604.16542cs.CRcs.CL2026-04被引 1

针对本地语言特点优化大模型安全防护,提升实际应用效果。

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

  • 基于台湾语境定制数据集,优化安全防护模型
  • F1提升0.289,误报率降低94.9%
  • 为地区性AI安全标准提供可落地的参考

安全防护机制已成为AI安全研究热点,旨在确保大语言模型(LLMs)行为得当。然而现有研究缺乏对语言与文化差异的考量,导致报告性能与实际应用效果之间存在差距。本文提出一种方法,通过利用针对特定语言特征的精选数据集,优化目标语言环境下的防护模型,以台湾语言环境为例,展示本地化部署的挑战。所提出的方案生成了TWGuard——一个针对语言上下文优化的安全防护模型,在性能上显著优于基础模型(F1提升0.289),在实际应用中更大幅降低误报率(-0.037,减少94.9%)。该工作为区域性社区建立基于自身语言背景的AI安全标准奠定基础,进一步验证了依赖主导语言标准的局限性。

原文摘要 · Abstract (English)

Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). However, existing research lacks consideration of nuances across linguistic and cultural contexts, resulting in a gap between reported performance and in-the-wild effectiveness. To address this issue, this paper proposes an approach to optimize guardrail models for a designated linguistic context by leveraging a curated dataset tailored to local linguistic characteristics, targeting the Taiwan linguistic context as a representative example of localized deployment challenges. The proposed approach yields TWGuard, a linguistic context-optimized guardrail model that achieves a huge gain (+0.289 in F1) compared to the foundation model and significantly outperforms the strongest baseline in practical use (-0.037 in false positive rate, a 94.9\% reduction). Together, this work lays a foundation for regional communities to establish AI safety standards grounded in their own linguistic contexts, rather than accepting boundaries imposed by dominant languages. The inadequacy of the latter is reconfirmed by our findings.

大模型安全本地化语言差异防护机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。