通过实时记录复现O1模型的过程,探索更开放透明的AI研究新范式。
O1 Replication Journey: A Strategic Progress Report -- Part 1
- 提出'旅程学习'范式,让模型学会完整探索过程而非仅找捷径。
- 仅用327个样本训练,数学推理能力比传统方法高出8%以上。
- 全程公开失败与进展,适合关注AI可解释性与开放科研的研究者。
本文介绍了一种开创性的智能研究方法——我们的O1复现之旅。面对OpenAI发布突破性O1模型的背景,我们启动一项透明、实时的复现探索,重新构想人工智能研究的开展与传播方式。该方法应对现代研究中团队封闭、信息延迟及贡献难被认可等关键问题。通过全面、实时地记录复现过程(含成功与失败),我们推动开放科学,加速集体进步,并为AI驱动的科学发现奠定基础。本报告不同于传统论文,提供持续更新、全过程透明和社区互动。技术上,提出‘旅程学习’范式,鼓励模型学习包括试错、反思和回溯在内的完整探索过程。在仅327个训练样本且无额外技巧的情况下,该方法在MATH数据集上性能超越传统监督学习超8%,展现出巨大潜力。我们认为这是已成功解码的O1核心技术。相关技术假设、认知探索图谱、自研工具等资源已公开于https://github.com/GAIR-NLP/O1-Journey。
原文摘要 · Abstract (English)
This paper introduces a pioneering approach to artificial intelligence research, embodied in our O1 Replication Journey. In response to the announcement of OpenAI's groundbreaking O1 model, we embark on a transparent, real-time exploration to replicate its capabilities while reimagining the process of conducting and communicating AI research. Our methodology addresses critical challenges in modern AI research, including the insularity of prolonged team-based projects, delayed information sharing, and the lack of recognition for diverse contributions. By providing comprehensive, real-time documentation of our replication efforts, including both successes and failures, we aim to foster open science, accelerate collective advancement, and lay the groundwork for AI-driven scientific discovery. Our research progress report diverges significantly from traditional research papers, offering continuous updates, full process transparency, and active community engagement throughout the research journey. Technologically, we proposed the journey learning paradigm, which encourages models to learn not just shortcuts, but the complete exploration process, including trial and error, reflection, and backtracking. With only 327 training samples and without any additional tricks, journey learning outperformed conventional supervised learning by over 8\% on the MATH dataset, demonstrating its extremely powerful potential. We believe this to be the most crucial component of O1 technology that we have successfully decoded. We share valuable resources including technical hypotheses and insights, cognitive exploration maps, custom-developed tools, etc at https://github.com/GAIR-NLP/O1-Journey.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。