用人类引导激发大模型解难题,成功攻克国际数学奥赛压轴题
Vibe Reasoning: Eliciting Frontier AI Mathematical Capabilities -- A Case Study on IMO 2025 Problem 6
- 通过通用元提示与智能体协作,唤醒大模型隐藏的数学能力
- 结合GPT-5与Gemini 3 Pro,算出正确答案2112并生成严谨证明
- 轻量人工干预即可释放前沿模型潜力,适合复杂推理任务研究者
我们提出Vibe Reasoning,一种人机协同解决复杂数学问题的新范式。核心观点是:前沿AI模型已具备解题所需知识,但缺乏如何、何时应用的意识。该方法通过通用元提示、智能体锚定和模型协同,将模型的潜在能力转化为实际解题能力。以IMO 2025第6题(组合优化题)为例,自主系统曾公开失败。我们的方案融合GPT-5的探索能力与Gemini 3 Pro的证明优势,采用带代码执行和文件记忆的智能体工作流,最终得出正确答案2112并完成严格证明。经多次迭代,发现智能体锚定与模型协同至关重要,而人类提示从具体线索逐步演变为通用可迁移的元提示。我们分析了模型自主失败的原因,阐明各组件如何应对特定失效模式,并提炼出有效Vibe Reasoning的原则。结果表明,轻量级人类引导可充分释放前沿模型的数学推理潜能。本工作仍在持续中,正开发自动化框架并开展更广范围评估,以验证该范式的普适性与有效性。
原文摘要 · Abstract (English)
We introduce Vibe Reasoning, a human-AI collaborative paradigm for solving complex mathematical problems. Our key insight is that frontier AI models already possess the knowledge required to solve challenging problems -- they simply do not know how, what, or when to apply it. Vibe Reasoning transforms AI's latent potential into manifested capability through generic meta-prompts, agentic grounding, and model orchestration. We demonstrate this paradigm through IMO 2025 Problem 6, a combinatorial optimization problem where autonomous AI systems publicly reported failures. Our solution combined GPT-5's exploratory capabilities with Gemini 3 Pro's proof strengths, leveraging agentic workflows with Python code execution and file-based memory, to derive both the correct answer (2112) and a rigorous mathematical proof. Through iterative refinement across multiple attempts, we discovered the necessity of agentic grounding and model orchestration, while human prompts evolved from problem-specific hints to generic, transferable meta-prompts. We analyze why capable AI fails autonomously, how each component addresses specific failure modes, and extract principles for effective vibe reasoning. Our findings suggest that lightweight human guidance can unlock frontier models' mathematical reasoning potential. This is ongoing work; we are developing automated frameworks and conducting broader evaluations to further validate Vibe Reasoning's generality and effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。