arXiv:2411.14405cs.CL2024-11被引 155

让大模型在无标准答案的开放问题上也能逻辑推理

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions

  • 用思维链+蒙特卡洛树搜索增强推理能力
  • 在数学物理编程外拓展至开放性问题解决
  • 适合需要创造性思考的复杂现实任务

当前OpenAI o1激发了对大型推理模型(LRM)的广泛关注。基于此趋势,Marco-o1不仅关注数学、物理、编程等有标准答案的领域——这些领域适合强化学习——更重视开放性解决方案。我们旨在回答:'o1模型能否有效泛化到缺乏明确标准、奖励难以量化的广阔领域?' Marco-o1采用思维链(CoT)微调、蒙特卡洛树搜索(MCTS)、反思机制及创新推理策略,专为复杂现实问题求解任务优化。

原文摘要 · Abstract (English)

Currently OpenAI o1 sparks a surge of interest in the study of large reasoning models (LRM). Building on this momentum, Marco-o1 not only focuses on disciplines with standard answers, such as mathematics, physics, and coding -- which are well-suited for reinforcement learning (RL) -- but also places greater emphasis on open-ended resolutions. We aim to address the question: ''Can the o1 model effectively generalize to broader domains where clear standards are absent and rewards are challenging to quantify?'' Marco-o1 is powered by Chain-of-Thought (CoT) fine-tuning, Monte Carlo Tree Search (MCTS), reflection mechanisms, and innovative reasoning strategies -- optimized for complex real-world problem-solving tasks.

大模型推理开放问题思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。