用可解释特征调控大模型,让预测更依赖历史而非未来信息。
Forecasting With LLMs: Improved Generalization Through Feature Steering

- 通过稀疏自编码器挖掘模型内部时间感知与前瞻偏差特征
- 增强时间感知特征后,前瞻偏差下降但推理能力保持不变
- 适合关注模型可解释性与公平预测的研究者
成功的预测需要识别历史与未来状态之间的可泛化模式。我们将在多种预测任务中应用大语言模型(LLMs),并使用稀疏自编码器分析其内部状态,以判断模型是依赖特定时间的知识,还是可泛化的模式。分析发现,存在与时间感知和前瞻偏差相关的特征。随后在全新领域中对这些特征进行干预。结果表明,增强时间感知特征能显著降低预测提示中的前瞻偏差,同时保持通用推理性能;而操纵候选的前瞻偏差特征则无明显效果。这表明可解释的时间特征可被用于因果性地引导模型转向更基于历史的推理。
原文摘要 · Abstract (English)
Successful forecasting involves identifying patterns between historical and future states of the world which generalize to future observations. We apply LLMs to a variety of forecasting tasks and inspect their internal states using sparse autoencoders to understand whether they appear to rely on time-specific pieces of knowledge versus generalizable patterns. Our analyses identify features associated with both time-aware reasoning and look-ahead-biased reasoning. We then apply the LLMs to an entirely different domain and intervene on these features. We find that amplifying time-awareness features substantially reduces look-ahead bias on forecasting prompts while preserving general reasoning performance. In contrast, steering the candidate look-ahead-bias features does not produce an effect. These results suggest that interpretable temporal features can be used to causally shift LLMs toward more historically grounded reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。