小模型通过分阶段强化学习,实现从理解过去到创造未来的完整时间推理能力。
Time-R1: Towards Comprehensive Temporal Reasoning in LLMs
- 分三阶段强化学习:先学历史事件逻辑,再预测未知未来,最后无微调生成创意场景。
- 30亿参数模型在预测和创作任务上超越6710亿参数的顶尖大模型。
- 适合对时间推理、可解释性生成感兴趣的开发者与研究者。
大型语言模型虽具强大能力,但在时间智能方面表现薄弱,难以整合过去推理与未来预测及创造。现有方法多聚焦单一时间技能,泛化能力差,尤其面对知识截止后的事件或需要创意展望的任务时表现不佳。为此,我们提出Time-R1,首个赋予中等规模(30亿参数)模型全面时间能力的框架——理解、预测与创造性生成。该框架采用新颖的三阶段发展路径:前两阶段构成由精心设计的动态规则奖励系统驱动的强化学习课程,逐步构建(1)基于历史数据的基础时间理解与事件时序映射,(2)超出知识截止时间的未来事件预测能力;最终(3)实现无需微调的创造性未来场景生成,具备卓越泛化能力。实验表明,Time-R1在极具挑战性的未来事件预测与创意场景生成基准上,显著优于超过200倍更大的模型,包括6710亿参数的SOTA模型DeepSeek-R1。本工作证明,经过精心设计的渐进式强化学习微调,使小型高效模型也能达到优越的时间推理性能,为真正具备时间感知的AI提供可行且可扩展的路径。为促进后续研究,我们还发布了基于十年新闻数据构建的大型多任务时间推理数据集Time-Bench,以及一系列Time-R1检查点。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate impressive capabilities but lack robust temporal intelligence, struggling to integrate reasoning about the past with predictions and plausible generations of the future. Meanwhile, existing methods typically target isolated temporal skills, such as question answering about past events or basic forecasting, and exhibit poor generalization, particularly when dealing with events beyond their knowledge cutoff or requiring creative foresight. To address these limitations, we introduce \textit{Time-R1}, the first framework to endow a moderate-sized (3B-parameter) LLM with comprehensive temporal abilities: understanding, prediction, and creative generation. Our approach features a novel three-stage development path; the first two constitute a \textit{reinforcement learning (RL) curriculum} driven by a meticulously designed dynamic rule-based reward system. This framework progressively builds (1) foundational temporal understanding and logical event-time mappings from historical data, (2) future event prediction skills for events beyond its knowledge cutoff, and finally (3) enables remarkable generalization to creative future scenario generation without any fine-tuning. Strikingly, experiments demonstrate that Time-R1 outperforms models over 200 times larger, including the state-of-the-art 671B DeepSeek-R1, on highly challenging future event prediction and creative scenario generation benchmarks. This work provides strong evidence that thoughtfully engineered, progressive RL fine-tuning allows smaller, efficient models to achieve superior temporal performance, offering a practical and scalable path towards truly time-aware AI. To foster further research, we also release \textit{Time-Bench}, a large-scale multi-task temporal reasoning dataset derived from 10 years of news data, and our series of \textit{Time-R1} checkpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。