语言模型用对数方式编码数值,但不等于能像人一样判断大小。
Weber's Law in Transformer Magnitude Representations: Efficient Coding, Representational Geometry, and Psychophysical Laws in Language Models
- 通过心理物理方法发现模型数值表征呈对数压缩几何。
- 96个模型-领域-层组合中对数几何相关性达0.68~0.96,线性几何不成立。
- 早期层负责数值处理,后期层虽有对数结构却不参与行为决策。
Transformer语言模型如何表示数量?现有研究存在分歧:有的发现对数间距,有的认为线性编码,还有人提出按位圆周表示。本文运用心理物理学的分析工具,在三个7-90亿参数的指令微调模型(涵盖Llama、Mistral、Qwen三类架构)的三个数量域中,通过表征相似性分析、行为辨别、精度梯度和因果干预四种方法进行验证。结果发现:第一,表征几何始终为对数压缩,与韦伯定律差异矩阵的相关系数在0.68至0.96之间,线性几何从未被偏好;第二,该几何与行为表现无关:一个模型表现出人类级韦伯分数(WF=0.20),另一个则无,且两者在时间与空间辨别任务中均表现随机;第三,因果干预显示层级分离:早期层在数量处理中功能特异性达4.1倍,而后期层尽管几何最强,其因果贡献仅1.2倍。语料库分析确认高效编码前提(α=0.77)。结果表明,训练数据统计足以生成对数压缩表征,但几何本身不足以保证行为能力。
原文摘要 · Abstract (English)
How do transformer language models represent magnitude? Recent work disagrees: some find logarithmic spacing, others linear encoding, others per-digit circular representations. We apply the formal tools of psychophysics to resolve this. Using four converging paradigms (representational similarity analysis, behavioural discrimination, precision gradients, causal intervention) across three magnitude domains in three 7-9B instruction-tuned models spanning three architecture families (Llama, Mistral, Qwen), we report three findings. First, representational geometry is consistently log-compressive: RSA correlations with a Weber-law dissimilarity matrix ranged from .68 to .96 across all 96 model-domain-layer cells, with linear geometry never preferred. Second, this geometry is dissociated from behaviour: one model produces a human-range Weber fraction (WF = 0.20) while the other does not, and both models perform at chance on temporal and spatial discrimination despite possessing logarithmic geometry. Third, causal intervention reveals a layer dissociation: early layers are functionally implicated in magnitude processing (4.1x specificity) while later layers where geometry is strongest are not causally engaged (1.2x). Corpus analysis confirms the efficient coding precondition (alpha = 0.77). These results suggest that training data statistics alone are sufficient to produce log-compressive magnitude geometry, but geometry alone does not guarantee behavioural competence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。