arXiv:2607.07918cs.LGcs.AI2026-07

用66个中性人格语句实现模型安全对齐,防攻不降效。

Efficient Safety Alignment of Language Models via Latent Personality Traits

论文配图:Efficient Safety Alignment of Language Models via Latent Personality Traits
图 1 · 摘自论文原文
  • 基于人格心理学语句做对抗训练,隐式约束越狱攻击空间。
  • 在HarmBench上攻击成功率接近零,且标准任务性能无损。
  • 训练仅需单卡几分钟,样本量仅为传统方法的1/75,适合快速部署。

当前大模型安全方法易受对抗攻击,亟需更鲁棒的替代方案。已有研究表明,潜在对抗训练(LAT)效果显著,但会损害模型效用且需大量有害提示数据。本文提出潜质人格对齐(LPA),仅使用66条来自心理测量学文献的中性人格陈述进行对抗训练,替代显式拒绝有害内容。我们假设人格锚定表征与避害行为共享潜在结构,因此对抗稳定这些表征可隐式限制越狱攻击所利用的子空间。LPA在HarmBench上对直接请求及五种越狱方法均实现近乎零的攻击成功率,训练过程中未接触任何有害内容,且标准基准性能无下降。整个训练过程轻量高效,单卡运行仅需数分钟,样本量仅为标准LAT的1/75。大量消融实验验证了该方法的鲁棒性、效率与泛化能力。

原文摘要 · Abstract (English)

Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT) is among the most effective defenses, but can degrade utility and requires training on large datasets of harmful prompts. We introduce Latent Personality Alignment (LPA), which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature. We hypothesize that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks. LPA achieves near-zero attack success rates on HarmBench across direct requests and five jailbreak methods, despite never seeing harmful content during training and no loss of performance on standard benchmarks. Moreover, the training process is lightweight; the entire procedure completes in minutes on a single GPU and uses 75x fewer examples than standard LAT. Extensive ablations demonstrate the robustness, efficiency, and generalization of our method.

安全对齐对抗训练轻量化人格建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。