揭示大模型如何内部表征时间偏好,并可操控其长远决策倾向。
Temporal Preference Concepts and their Functions in a Large Language Model

- 通过梯度归因与激活修补定位时间偏好子图,发现中上层节点编码时间维度。
- 模型对未来奖励折扣率比人类低数倍,且在不同情境下不稳定。
- 可通过方向向量调节模型的时间偏好,适合对长期规划有要求的研究者。
大型语言模型(LLMs)越来越多地用于需要权衡短期收益与长期后果的决策任务,但对其内部如何表示或解决此类权衡仍知之甚少。本文在精简版LLM(Qwen3-4B-Instruct-2507)中,通过因果定位识别出一个与时间偏好相关的底层子图,利用梯度归因和激活修补的汇聚证据,确定了中上层节点的作用。我们发现时间跨度的几何结构编码在预期定位层的残差流中。行为分析显示,未经干预的模型对未来奖励的折扣程度远低于人类(数倍),且该偏好在不同上下文中表现不稳定,这表明应主动控制而非依赖训练隐含偏好。最后,我们发现了提示性证据:引导向量可改变时间偏好。本研究展示了机制可解释性如何推动对大模型规划与推理的可靠控制。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly being deployed to make decisions that require trading off near-term gains against long-term consequences, yet little is known about how they internally represent or resolve these tradeoffs. In this work, we causally localize an underlying subgraph for temporal preference in a distilled LLM (Qwen3-4B-Instruct-2507), identifying mid-to-upper-layer nodes through converging evidence from gradient-based attribution and activation patching. We find that the geometry of time horizon is encoded in the residual stream at the expected localized layers. A behavioral analysis reveals that unintervened LLMs discount the future several times less steeply than humans, yet this preference is unstable across contexts, motivating explicit control rather than implicit reliance on training. Finally, we find suggestive evidence that steering vectors can shift temporal preference. Our work demonstrates how mechanistic interpretability can bring us closer to reliable control over how LLMs plan and reason
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。