arXiv:2602.19101cs.CLcs.AI2026-02

发现大模型对道德、语法、经济价值混淆,修复后表现更准

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

  • 通过探针分析模型行为与激活向量,发现三类价值混杂
  • 语法和经济估值受道德影响过度,偏离人类基准
  • 仅移除道德相关激活向量即可显著修复价值混淆

大型语言模型(LLMs)的价值对齐需实证测量其实际价值表征。人类价值表征的一个特征是能区分不同类型的善:道德、语法和经济。我们探究了LLMs是否也具备这种区分能力。通过探测模型行为、嵌入表示和残差流激活,发现普遍存在价值纠缠现象——这三类价值表征被混淆。具体而言,相较于人类标准,语法与经济估值均过度受道德价值影响。通过选择性消融与道德相关的激活向量,该混淆现象得以修复。

原文摘要 · Abstract (English)

Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Among the characteristics of value representation in humans is that they distinguish among value of different kinds. We investigate whether LLMs likewise distinguish three different kinds of good: moral, grammatical, and economic. By probing model behavior, embeddings, and residual stream activations, we report pervasive cases of value entanglement: a conflation between these distinct representations of value. Specifically, both grammatical and economic valuation was found to be overly influenced by moral value, relative to human norms. This conflation was repaired by selective ablation of the activation vectors associated with morality.

大模型对齐价值表征模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。