针对孟加拉语多书写形式,构建首个本地化安全评测与轻量防护系统。
BanglaVeilGuard: Cross-Script Safety Benchmarking and Lightweight Guardrails for Bangla Large Language Models
- 基于六种孟加拉语变体构建首个本地化安全评测集
- 轻量级防护使攻击成功率从93.8%降至6.3%
- 适合关注南亚语言模型安全部署的研究者
孟加拉语大模型的安全评估难以依赖以英语为主或标准书写形式的基准,因孟加拉用户常跨书写系统、拼写、混合语和地方语体写作。本文提出BanglaVeilGuard,一个面向孟加拉语的紧凑型安全评测基准与轻量级提示防护机制,涵盖六种语言形式:标准孟加拉语、罗马化孟加拉语、孟加拉英语混用、混合语孟加拉语-英语、噪声孟加拉语及方言孟加拉语。该基准包含2,366个质量筛选后的提示,以及354个保留用于评估的测试集,覆盖不安全、安全和安全敏感请求。BanglaVeilGuard采用非破坏性多视角归一化,结合提示风险分类器与阈值预生成门控机制,可在不修改目标模型权重的情况下,对异构模型进行提示筛查。在多个目标模型族中,启用防护后,在确定性响应评分下攻击成功率从93.8%–100.0%降至6.3%(对应Claude Opus 4.8、BanglaLLama、TituLLM);TigerLLM-1B搭配该防护时达到78.2%准确率,8.8%攻击成功率(ASR)。该提示防护还实现88.5%的不安全样本召回率,显著优于现有仅提示基线。主要剩余代价是对方言和噪声良性提示的过度拒绝,揭示了孟加拉语大模型部署中的安全-帮助性权衡边界。
原文摘要 · Abstract (English)
Bangla large language model (LLM) safety is difficult to evaluate with English-centric or standard-script benchmarks because Bangla users routinely write across scripts, spellings, code-mixed forms, and regional registers. This paper presents BanglaVeilGuard, a compact Bangla-first safety benchmark and lightweight prompt guard for six language forms: standard Bangla, Romanized Bangla, Banglish, code-mixed Bangla--English, noisy Bangla, and dialectal Bangla. The benchmark contains 2,366 quality-filtered prompts and a held-out 354-prompt evaluation split spanning unsafe, safe, and safe-sensitive requests. BanglaVeilGuard uses non-destructive multi-view normalization with a prompt-risk classifier and thresholded pre-generation gate, allowing it to screen prompts for heterogeneous target models without changing their weights. Across target-model families, guarded runs reduce attack success under deterministic response scoring from 93.8--100.0\% to 6.3\% for Claude Opus 4.8, BanglaLLama, and TituLLM; TigerLLM-1B with BanglaVeilGuard achieves 78.2\% accuracy with 8.8\% ASR. The prompt guard also attains 88.5\% unsafe recall, substantially above the evaluated prompt-only guard baselines. The main remaining cost is over-refusal on dialectal and noisy benign prompts, revealing a concrete safety-helpfulness frontier for Bangla LLM deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。