arXiv:2606.15517cs.CL2026-06

让大模型学会在安全前提下更贴心地回应敏感问题。

SHARD: Safe and Helpful Alignment via Self-Reframing Distillation

论文配图:SHARD: Safe and Helpful Alignment via Self-Reframing Distillation
图 1 · 摘自论文原文
  • 用哲学原则重写敏感问题,挖掘潜在合理意图
  • 自重构回答使帮助性提升,同时保持安全性
  • 无需大模型教师,小模型也能学会安全又贴心的回应

大语言模型在面对敏感提示时常表现不佳:要么直接拒绝,要么给出通用的安全套话,或无法满足用户可安全回答的信息需求。本文提出SHARD,一种自重构蒸馏方法,通过哲学指导重写敏感提示以揭示善意意图,再将原始回答重构为安全且更有帮助的形式,并以此自生成数据微调模型。在DNA和LINGUASAFE英文子集上的实验表明,SHARD在多数模型族中提升了帮助性,同时维持了安全性。该方法在性能上仍能媲美使用更大教师模型的蒸馏,说明模型能从自身生成的反馈中内化安全且有帮助的行为。警告:本文包含可能令人不适或有害的内容。

原文摘要 · Abstract (English)

Large language models often struggle with sensitive prompts. They may refuse outright, provide generic safety boilerplate, or fail to address the user's legitimate informational needs that can be answered safely. We introduce SHARD, a self-reframing distillation method to improve safe-helpfulness. It first rewrites sensitive prompts to surface benign intent using philosophical guidelines, then reframes its original responses into safe, more helpful ones, and finally fine-tunes the model on its self-reframed responses. Across DNA and the English subset of LINGUASAFE, SHARD improves helpfulness for most model families while preserving safety. It also remains competitive with distillation from a larger teacher model, suggesting that models can internalize safe and helpful behavior elicited from their own. Warning: This paper contains content that may be offensive or harmful.

大模型对齐安全对话自重构提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。