arXiv:2607.25554cs.AI2026-07

用时间截断机制提升大模型未来预测能力,避免历史信息泄露。

Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis

论文配图:Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis
图 1 · 摘自论文原文
  • 引入时间截断钩子,每轮强制限制历史数据范围
  • 生成数据效率提升,高质量数据占比提高37%
  • 适合需要长期趋势推理的预测任务研究者

未来事件预测具有广泛社会影响但依然困难。当前最优方法依赖外部代理框架,一旦移除便失效。尽管近期工具集成推理(TIR)已内化多跳事实检索能力,但预测还需对历史趋势与动态变化进行时间维度的搜索与推理。核心瓶颈在于数据:历史查询导致时间泄漏,使预测退化为单纯检索。以往方法或冻结信息收集,或依赖拒绝采样或未解决的新查询,造成大量数据浪费,降低合成效率。本文提出时间截断钩子,在每一步强制设定时间截止点,实现类TIR的采样,减少时间泄漏及对拒绝采样或未解查询的依赖,提升采样效率。我们进一步构建大规模语料库与基于过程的评估指标,验证该钩子自然扩大了时间搜索广度,显著提高高质量数据比例,进一步增强效率并降低对复杂评判标准的依赖。蒸馏实验表明,学生模型在钩子干预数据上训练后表现最佳,证明钩子辅助的模型演进可将更高质量的时间搜索与推理数据转化为参数层面的性能提升。

原文摘要 · Abstract (English)

Future event prediction carries broad social impact yet remains challenging. SOTA approaches augment LLMs with external agent frameworks whose predictive capability vanishes once the harness is removed. While recent Tool-Integrated Reasoning (TIR) internalizes deep search for multi-hop retrieval of facts, forecasting further demands temporal search and reasoning over historical trends and dynamic shifts. The key obstacle is data: historical queries induce temporal leakage that degrades forecasting into retrieval. Prior works either freeze information gathering with static observations, or rely on rejection sampling or unresolved fresh queries that discard vast amounts of data, degrading synthesis efficiency. We propose a time-truncation harness that enforces a temporal cut-off at every turn, enabling TIR-style sampling from historical events, reducing temporal leakage and reliance of rejection sampling or unsolved queries, increasing the sampling efficiency. We further build a large-scale corpus and a process-based metric and show that our harness naturally induces a broader temporal breadth of search and raises the proportion of high-quality data, further increasing the efficiency and reducing the reliance on complex rubrics. Distillation experiments show that students trained on harness-intervened data achieve the best performance, demonstrating harness-assisted model evolving that turns higher quality temporal search and reasoning data into a parametric advancement of the students.

未来预测时间推理大模型蒸馏数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。