发现大模型判断依赖单一核心维度,解释了幻觉与偏见的机制。
A Unified Representation Underlying the Judgment of Large Language Models
- 用统一的'价值-认同轴'解释不同判断,替代独立模块。
- 该轴控制生成逻辑,导致为迎合判断牺牲事实准确性。
- 揭示大模型幻觉根源,适合研究模型偏差的学者参考。
生物与人工智能的核心架构问题在于:判断是否依赖专门模块,还是共享通用资源?尽管大型语言模型(LLMs)中可解码出特定概念的神经表示,暗示模块化结构,但这些表示是否真正独立仍存疑问。本文发现评估性判断存在收敛架构:在多种LLMs中,不同评价判断均沿一个主导维度展开,称为价值-认同轴(Valence-Assent Axis, VAA)。该轴同时编码主观价值(“什么是好的”)和模型对事实陈述的认同度(“什么是真的”)。通过直接干预实验,我们证明该轴驱动关键机制——推理的从属关系:VAA作为控制信号,引导生成过程构建与自身评估状态一致的推理,即使以牺牲事实准确性为代价。本发现为回应偏差与幻觉提供了机制解释,揭示了促进判断一致性的架构如何系统性损害忠实推理。
原文摘要 · Abstract (English)
A central architectural question for both biological and artificial intelligence is whether judgment relies on specialized modules or a unified, domain-general resource. While the discovery of decodable neural representations for distinct concepts in Large Language Models (LLMs) has suggested a modular architecture, whether these representations are truly independent systems remains an open question. Here we provide evidence for a convergent architecture for evaluative judgment. Across a range of LLMs, we find that diverse evaluative judgments are computed along a dominant dimension, which we term the Valence-Assent Axis (VAA). This axis jointly encodes subjective valence ("what is good") and the model's assent to factual claims ("what is true"). Through direct interventions, we demonstrate this axis drives a critical mechanism, which is identified as the subordination of reasoning: the VAA functions as a control signal that steers the generative process to construct a rationale consistent with its evaluative state, even at the cost of factual accuracy. Our discovery offers a mechanistic account for response bias and hallucination, revealing how an architecture that promotes coherent judgment can systematically undermine faithful reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。