针对台湾闽南语安全风险,构建专用评测基准与防护模型。
Taiwan Safety Benchmark and Breeze Guard: Toward Trustworthy AI for Taiwanese Mandarin
- 基于台湾文化语料训练安全模型,避免通用模型的地域盲区。
- 在本地化风险检测上,模型表现优于主流8B模型0.17分(最高+0.66)。
- 适合关注台湾本土化AI安全的开发者与政策制定者使用。
全球安全模型在通用基准上表现良好,但其训练数据缺乏台湾闽南语的文化与语言细节,导致对本地化风险(如金融诈骗、嵌入式仇恨言论、虚假信息)存在系统性盲区。为此,我们提出TS-Bench(台湾安全基准),一个涵盖400个经人工标注提示的标准化评估套件,覆盖金融欺诈、医疗谣言、社会歧视和政治操控等关键领域。同时,推出基于先前发布的通用型台湾闽南语大模型Breeze 2微调而成的8B安全模型Breeze Guard,该模型通过大规模人工验证合成数据进行监督微调,聚焦台湾特定危害。核心假设是:有效的安全检测需依托基座模型已有的文化语境;仅靠安全微调无法从零构建社会语言知识。实证表明,Breeze Guard在TS-Bench上显著优于领先的8B通用安全模型Granite Guardian 3.3(整体F1提升0.17),尤其在高上下文场景中表现突出,如诈骗类(+0.66 F1)、金融违规类(+0.43 F1)。尽管在英文主导的基准(ToxicChat、AegisSafetyTest)上略有下降,但这是区域性专精模型的合理权衡。二者共同为台湾可信AI部署奠定新基础。
原文摘要 · Abstract (English)
Global safety models exhibit strong performance across widely used benchmarks, yet their training data rarely captures the cultural and linguistic nuances of Taiwanese Mandarin. This limitation results in systematic blind spots when interpreting region-specific risks such as localized financial scams, culturally embedded hate speech, and misinformation patterns. To address these gaps, we introduce TS-Bench (Taiwan Safety Benchmark), a standardized evaluation suite for assessing safety performance in Taiwanese Mandarin. TS-Bench contains 400 human-curated prompts spanning critical domains including financial fraud, medical misinformation, social discrimination, and political manipulation. In parallel, we present Breeze Guard, an 8B safety model derived from Breeze 2, our previously released general-purpose Taiwanese Mandarin LLM with strong cultural grounding from its original pre-training corpus. Breeze Guard is obtained through supervised fine-tuning on a large-scale, human-verified synthesized dataset targeting Taiwan-specific harms. Our central hypothesis is that effective safety detection requires the cultural grounding already present in the base model; safety fine-tuning alone is insufficient to introduce new socio linguistic knowledge from scratch. Empirically, Breeze Guard significantly outperforms the leading 8B general-purpose safety model, Granite Guardian 3.3, on TS-Bench (+0.17 overall F1), with particularly large gains in high-context categories such as scam (+0.66 F1) and financial malpractice (+0.43 F1). While the model shows slightly lower performance on English-centric benchmarks (ToxicChat, AegisSafetyTest), this tradeoff is expected for a regionally specialized safety model optimized for Taiwanese Mandarin. Together, Breeze Guard and TS-Bench establish a new foundation for trustworthy AI deployment in Taiwan.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。