针对阿拉伯语模型的安全对齐,提出按需拒绝策略并验证其有效性
Arabic Safety Alignment as Selective Refusal: An Empirical Study of SFT, DPO, and Guard Calibration

- 采用有害提示拒绝率与良性提示拒绝率双指标评估模型安全对齐效果
- 精选SFT配置使有害拒绝率达90%~93%,良性拒绝率控制在14%~23%
- 不同模型需个性化调整,无统一最优对齐方法
阿拉伯语大模型需在拒绝有害提示的同时避免过度拒绝良性或敏感提示,但单一拒绝率掩盖了这一权衡。我们通过良性拒绝率B和有害提示拒绝率H来评估,其中H衡量拒绝行为而非有害内容的顺从。在五个具备阿拉伯语能力的模型上,对完整人工撰写的人类数据集AraSafe进行130次实验,仅拒绝的监督微调(SFT)导致全拒倾向;而精选的混合SFT配置可在保持B=14%~23%的前提下实现H=90%~93%;其中四种配置在全部三轮中超过H=90%目标,Fanar则在两轮中达标。直接偏好优化(DPO)与推理防护机制对不同模型的影响各异,非统一提升。在盲审的300条响应中,标注者二元拒绝一致性达89.0%(kappa=0.78);Qwen3Guard与Aya Expanse 32B分别达到88.7%与91.0%准确率,无显著差异。精选SFT在阿拉伯语变体(Arabizi)上提升了所有五模型的有害拒绝率,但均未达90%,表明从中东标准阿拉伯语迁移有限。总体结果支持模型特异性操作点选择:设定部署目标,并仅保留能提升该目标的干预措施。
原文摘要 · Abstract (English)
Arabic large language models must refuse harmful prompts without over-refusing benign or sensitive prompts, yet a single refusal rate hides this trade-off. We evaluate it using benign refusal B and harmful-prompt refusal H, where H measures refusal rather than harmful compliance. Across five Arabic-capable models and 130 runs on the full human-written AraSafe set, refusal-only supervised fine-tuning (SFT) collapses toward blanket refusal, whereas selected mixed-SFT configurations reach H = 90% to 93% at B = 14% to 23%; four selected configurations exceed the H = 90% target in all three runs, while Fanar does so in two of three. Direct Preference Optimization (DPO) and inference guards change B and H differently across models rather than acting as uniform upgrades. In a blinded 300-response audit, annotator binary-refusal agreement is 89.0% (kappa = 0.78); Qwen3Guard and Aya Expanse 32B reach 88.7% and 91.0% accuracy, respectively, with no conclusive paired difference. Selected SFT raises H on Arabizi for all five models, but none reaches 90%, showing only partial transfer from Modern Standard Arabic. Overall, the results support model-specific operating-point selection: set a deployment target and retain only interventions that improve it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。