arXiv:2609.04022cs.CLcs.AI2026-09

让大模型理解道德概念的内在结构,提升对抗攻击下的安全性。

Representational alignment yields generalizable safety in language models

  • 通过优化模型内部表征与人类道德判断的相似性,而非只对输出做监督。
  • 在25万条道德标注上测试,显著提升模型在对抗攻击下的鲁棒性。
  • 适用于不同规模模型,为安全可控的AI提供新思路,适合安全研究者参考。

大型语言模型(LLMs)的安全部署依赖于对齐。现有方法主要优化可观察的响应,但当相同有害意图以陌生或对抗形式出现时,模型仍易失效,而人类却能轻易识别。原型理论解释了这种适应性:人类概念围绕典型样本构建,新实例按其与原型的典型性程度分类。本文发现,当前23个LLMs中,道德概念的分类结构在模型中弱化,难以区分对立道德类别,也未保留类别内的精细典型性。这些缺陷在不同参数量和对齐阶段均存在。我们提出表示相似性优化(representational similarity optimization),直接对齐模型潜在表征与人类道德判断的分类结构,无需监督生成响应。在相同251,334条道德标注下,标准行为对齐虽提升了显式判断准确率,但未改变分类结构,且加剧对抗脆弱性;而重构道德分类结构虽对显式判断提升有限,却在多种基准和攻击策略下一致增强各规模模型的对抗鲁棒性。结果支持原型分类促进行为适应性的观点,并表明将此表征原则迁移至LLMs可实现对抗条件下的通用安全。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.

模型安全道德对齐表征学习对抗鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。