arXiv:2510.13166cs.CL2025-10被引 4

用进化方法优化大模型的科学推理过程,生成高质量训练数据

CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning

  • 通过多模型生成、知识增强与迭代进化,提升推理路径质量
  • 在多个科学推理基准上超越现有小模型性能,达到新SOTA
  • 适合需要高质量科学推理能力的轻量级模型开发者

尽管从大型语言模型(LLMs)中进行链式思维(CoT)蒸馏在通用推理任务中已证明有效,但在科学领域仍面临挑战——即使先进模型也常因复杂性和专业知识要求而产生错误或浅层推理。直接蒸馏这些有缺陷的输出会生成低质量训练数据,限制小型学生模型的表现。为此,我们提出CoT-Evo,一种基于进化的链式思维蒸馏框架。该框架首先从多个LLM思考者中构建多样化的推理轨迹,自动引入领域知识进行丰富,并通过新颖性驱动的选择、反思性重组与突变进行迭代优化。优化过程由一个评估答案正确性、连贯性和知识利用效率的适应度函数指导,最终生成专用于科学推理的高质量CoT数据集。我们使用该数据集微调紧凑模型,在多个科学推理基准上取得当前最优表现。本工作建立了一种可扩展的方法,能够从多样且易出错的LLMs中合成高保真科学推理数据。

原文摘要 · Abstract (English)

While chain-of-thought (CoT) distillation from advanced large language models (LLMs) has proven effective in general reasoning tasks, it struggles in scientific domains where even advanced models often produce incorrect or superficial reasoning due to high complexity and specialized knowledge requirements. Directly distilling from such flawed outputs results in low-quality training data and limits the performance of smaller student models. To overcome this, we propose CoT-Evo, an evolutionary CoT distillation framework. It begins by constructing a diverse pool of reasoning trajectories from multiple LLM thinkers, enriches them with automatically retrieved domain knowledge, and iteratively refines the trajectories using novelty-driven selection, reflective recombination and mutation. The refinement is guided by a fitness function that evaluates answer correctness, coherence, and effective knowledge utilization. This results in a high-quality CoT dataset tailored for scientific reasoning. We employ this evolved dataset to fine-tune a compact model, which achieves state-of-the-art performance on scientific reasoning benchmarks. Our work establishes a scalable approach to synthesizing high-fidelity scientific reasoning data from diverse and fallible LLMs.

链式思维科学推理模型蒸馏进化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。