arXiv:2510.11677cs.LGq-fin.GN2025-10被引 3

构建时间一致的指令微调模型,消除未来信息泄露问题。

Instruction Tuning Chronologically Consistent Language Models

  • 基于截止日期前数据训练,确保时间顺序无泄露。
  • 提供可复现的开放模型权重和对话式交互界面。
  • 适合需要严格时间约束的预测研究,如历史分析与趋势推演。

我们提出一系列时间一致的指令微调大语言模型,以消除前瞻偏差。每个模型仅使用在明确知识截止日期前可用的数据进行训练,确保与截止日期后数据严格隔离。该框架具备(i)简洁的对话式聊天界面,(ii)完全开放且固定的模型权重,保证结果可复现,(iii)保守的预测准确率下界,剥离训练泄漏后仍能保留的可预测性部分。这些特性共同为研究人员提供了一种易用、无前瞻偏差的生成式AI工具,适用于广泛预测任务。

原文摘要 · Abstract (English)

We introduce a family of chronologically consistent, instruction-tuned large language models to eliminate lookahead bias. Each model is trained only on data available before a clearly defined knowledge-cutoff date, ensuring strict temporal separation from any post-cutoff data. The resulting framework offers (i) a simple, conversational chat interface, (ii) fully open, fixed model weights that guarantee replicability, and (iii) a conservative lower bound on forecast accuracy, isolating the share of predictability that survives once training leakage is removed. Together, these features provide researchers with an easy-to-use generative AI tool useful for a wide range of prediction tasks that is free of lookahead bias.

指令微调时间一致性预测建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。