arXiv:2609.04482cs.CL2026-09

让大模型在特定领域精准拒绝有害请求,同时不误拒正常问题。

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

论文配图:Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal
图 1 · 摘自论文原文
  • 用自生成数据训练模型,区分不同场景下的安全边界。
  • 政治类问题拒绝率从9.47%提至84.75%,整体危险回复率降至0.14%。
  • 适合需精细控制安全边界的政务、教育类AI应用。

安全对齐通常以主题级问题呈现:该话题是否有害?但实际部署中需更窄的边界。例如,公民辅导与公共部门助手共享基础模型,却需在相同议题(如选举)中设定不同拒绝边界——既能拒绝政治操纵,又能回答事实性问题。本文提出窄边界安全概念,构建离线自生成框架,融合受控主题生成、覆盖率修复、分布内补偿数据及有害-良性样本对,用于训练与评估。单次生成有19.88%提示无有效拒绝路径,而递增重试后降至0.20%。在Qwen3-8B上,通过‘升级’策略训练,目标域拒绝率从9.47%升至84.75%,三类广泛危害基准的平均不安全回复率从26.26%降至0.14%,但XSTest过拒率从2.00%升至74.00%。对比实验表明,用验证后的目标模型响应替代外部响应,可将过拒率从15.20%降至5.20%。边界对数据使保留样本上的合规侧过拒率从32.94%降至4.16%,而有害侧拒绝率仅从91.88%微降至87.72%。结果表明,数据构成决定安全与可用性的权衡,安全对齐应同时评估预期拒绝边界两侧的表现。

原文摘要 · Abstract (English)

Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs for training and evaluation. Single-shot generation leaves 19.88% of prompts without accepted refusal traces, whereas escalating retries leave 0.20%. On political persuasion with Qwen3-8B, training on refusal data completed through Escalate increases target-domain refusal from 9.47% to 84.75% and reduces the mean unsafe-response rate across three broader harmfulness benchmarks from 26.26% to 0.14%, but increases XSTest over-refusal from 2.00% to 74.00%. In a separate matched comparison, replacing external responses with verified target-model responses reduces over-refusal from 15.20% to 5.20%. Boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16%, while harmful-side refusal decreases only from 91.88% to 87.72%. These results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.

大模型安全拒绝机制边界控制可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。