通过多阶段批判性推理,提升多模态模型的科学解题能力。
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
- 设计全流程教育推理框架,融合动态反馈与自省机制。
- 在数学物理任务上,推理一致性和连贯性显著提升。
- 适合研究教育智能、科学推理与模型自我修正的学者。
多模态大语言模型在科学任务上表现仍不佳,尤其在需要多步可解释推理的任务中。其局限包括缺乏科学推理模式、多步推导缺乏全局连贯性,以及缺少反思性自我修正能力,导致在结构化科学场景中不可靠。我们提出EduFlow,首个覆盖教育科学推理全链条的端到端框架,包含数据选择、基于MCTS的轨迹构建、模型训练和输出优化。核心是EduPRM,一种过程感知奖励模型,能对推理步骤进行标签化和理由化批判。EduPRM通过课程学习,利用三种互补监督信号训练:MCTS引导的轨迹、注入错误的批判和师生对话,实现对多阶段问题求解的动态适应和推理中的迭代优化。我们进一步提出EduMCTS,一种面向教育推理的领域适配搜索框架,引入如自省机制等专门的启动动作,促进错误反思与修正,并利用EduPRM的细粒度反馈引导搜索至高质量推理轨迹。通过自一致性与拒绝采样,构建了EduMCTS-160K,一个大规模教育推理轨迹数据集。大量实验表明,EduFlow显著提升了推理的一致性与连贯性。代码、数据与模型将公开。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) still perform poorly on scientific tasks, particularly those requiring multi-step and interpretable reasoning. Their limitations include insufficient scientific reasoning patterns, lack of global coherence in multi-step inference, and the absence of reflective self-correction, making them unreliable in structured scientific contexts. We introduce EduFlow, the first end-to-end framework that covers the full pipeline of educational scientific reasoning, including data selection, MCTS-based trajectory construction, model training, and output optimization. At its core is EduPRM, a process-aware reward model that critiques reasoning steps with tags and justifications. EduPRM is trained via curriculum learning on three complementary supervision sources: MCTS-guided trajectories, error-injected critiques, and teacher-student dialogues, enabling dynamic adaptation to multi-stage problem solving and iterative refinement during inference. We further propose EduMCTS, a domain-adapted search framework that introduces bootstrapping actions specifically designed for educational reasoning, such as a self-reflection mechanism that promotes reflective error correction. It further leverages EduPRM's fine-grained feedback to guide the search toward higher-quality reasoning trajectories. By applying self-consistency and rejection sampling, we constructed EduMCTS-160K, a large-scale dataset of educational reasoning trajectories. Extensive experiments demonstrate that EduFlow enhances reasoning consistency and coherence. Code, data, and models will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。