arXiv:2605.08496cs.AI2026-05

用抽象人格特质训练模型,不用有害样本也能防攻击

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms

论文配图:Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
图 1 · 摘自论文原文
  • 不依赖具体有害语句,转而训练模型对人格特质的认同
  • 仅用不到100条特质描述,防御效果媲美15万+样本方法
  • 对未知攻击类型泛化能力强,误判率降低2.6倍

当前大语言模型的对抗鲁棒性方法需大量有害提示数据(数千至数十万条),但仍易受新型攻击和分布偏移影响。我们提出隐空间人格对齐(LPA),通过在抽象人格特质上训练模型,而非具体有害行为,实现高效防御。仅使用不足100条特质语句与隐式对抗训练,LPA在攻击成功率上达到训练集超15万条样本方法的水平,同时保持更高实用性。关键在于,LPA在六种危害基准测试中对未见攻击分布的泛化能力更强,误分类率比基线降低2.6倍,且训练过程从未接触过有害示例。结果表明,基于人格的对齐为低成本构建稳健防御提供了系统性路径。

原文摘要 · Abstract (English)

Current adversarial robustness methods for large language models require extensive datasets of harmful prompts (thousands to hundreds of thousands of examples), yet remain vulnerable to novel attack vectors and distributional shifts. We propose Latent Personality Alignment (LPA), a sample-efficient defense that achieves robustness by training models on abstract personality traits rather than specific harmful behaviors. Using fewer than 100 trait statements and latent adversarial training, LPA achieves comparable attack success rates to methods trained on 150k+ examples, while maintaining superior utility. Critically, LPA generalizes better to unseen attack distributions, reducing misclassification rates by 2.6x compared to baseline across six harm benchmarks -- without ever seeing harmful examples during training. Our results demonstrate that personality-based alignment offers a principled approach to building robust defenses with minimal cost.

模型安全对抗鲁棒性人格对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。