让大模型在特定领域精准拒绝有害请求,同时不误拒正常问题。
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

- 用自生成数据训练模型,区分不同场景下的安全边界。
- 政治类问题拒绝率从9.47%提至84.75%,整体危险回复率降至0.14%。
- 适合需精细控制安全边界的政务、教育类AI应用。
安全对齐通常以主题级问题呈现:该话题是否有害?但实际部署中需更窄的边界。例如,公民辅导与公共部门助手共享基础模型,却需在相同议题(如选举)中设定不同拒绝边界——既能拒绝政治操纵,又能回答事实性问题。本文提出窄边界安全概念,构建离线自生成框架,融合受控主题生成、覆盖率修复、分布内补偿数据及有害-良性样本对,用于训练与评估。单次生成有19.88%提示无有效拒绝路径,而递增重试后降至0.20%。在Qwen3-8B上,通过‘升级’策略训练,目标域拒绝率从9.47%升至84.75%,三类广泛危害基准的平均不安全回复率从26.26%降至0.14%,但XSTest过拒率从2.00%升至74.00%。对比实验表明,用验证后的目标模型响应替代外部响应,可将过拒率从15.20%降至5.20%。边界对数据使保留样本上的合规侧过拒率从32.94%降至4.16%,而有害侧拒绝率仅从91.88%微降至87.72%。结果表明,数据构成决定安全与可用性的权衡,安全对齐应同时评估预期拒绝边界两侧的表现。
原文摘要 · Abstract (English)
Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs for training and evaluation. Single-shot generation leaves 19.88% of prompts without accepted refusal traces, whereas escalating retries leave 0.20%. On political persuasion with Qwen3-8B, training on refusal data completed through Escalate increases target-domain refusal from 9.47% to 84.75% and reduces the mean unsafe-response rate across three broader harmfulness benchmarks from 26.26% to 0.14%, but increases XSTest over-refusal from 2.00% to 74.00%. In a separate matched comparison, replacing external responses with verified target-model responses reduces over-refusal from 15.20% to 5.20%. Boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16%, while harmful-side refusal decreases only from 91.88% to 87.72%. These results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。