大模型自发形成时间参照点,呈现人类特有的对时间的感知规律。
The Other Mind: How Language Models Exhibit Human Temporal Cognition
- 通过相似性任务发现模型自发建立主观时间参考点。
- 时间感知距离按对数压缩,符合韦伯-费希纳定律。
- 适合关注AI认知机制与人机对齐的研究者阅读。
随着大语言模型(LLMs)的发展,它们展现出一些未在训练数据中明确指定但与人类认知相似的模式。本研究聚焦于模型的时间认知能力,通过相似性判断任务发现,更大模型会自发建立主观时间参考点,并遵循韦伯-费希纳定律——感知时间距离随年份远离参考点而对数压缩。我们从神经元、表征和信息三个层面分析其机制:识别出一组时间偏好神经元,在主观参考点处激活最低,并采用与生物系统一致的对数编码;年份表征呈现层级结构,浅层为数值,深层演化为抽象时间方向;预训练嵌入分析显示,训练语料本身具有非线性的内在时间结构,为模型内部构建提供了基础。讨论中提出经验主义视角,认为模型认知是内部表征系统对世界的主观建构。这一观点暗示了可能涌现人类难以直觉预测的异质认知框架,指向以引导内部建构为核心的新型人工智能对齐方向。代码开源:https://TheOtherMind.github.io。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) continue to advance, they exhibit certain cognitive patterns similar to those of humans that are not directly specified in training data. This study investigates this phenomenon by focusing on temporal cognition in LLMs. Leveraging the similarity judgment task, we find that larger models spontaneously establish a subjective temporal reference point and adhere to the Weber-Fechner law, whereby the perceived distance logarithmically compresses as years recede from this reference point. To uncover the mechanisms behind this behavior, we conducted multiple analyses across neuronal, representational, and informational levels. We first identify a set of temporal-preferential neurons and find that this group exhibits minimal activation at the subjective reference point and implements a logarithmic coding scheme convergently found in biological systems. Probing representations of years reveals a hierarchical construction process, where years evolve from basic numerical values in shallow layers to abstract temporal orientation in deep layers. Finally, using pre-trained embedding models, we found that the training corpus itself possesses an inherent, non-linear temporal structure, which provides the raw material for the model's internal construction. In discussion, we propose an experientialist perspective for understanding these findings, where the LLMs' cognition is viewed as a subjective construction of the external world by its internal representational system. This nuanced perspective implies the potential emergence of alien cognitive frameworks that humans cannot intuitively predict, pointing toward a direction for AI alignment that focuses on guiding internal constructions. Our code is available at https://TheOtherMind.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。