arXiv:2507.23751cs.AIcs.CL2025-07被引 29

用思维链自动生成高质量合成数据,提升大模型推理与指令遵循能力。

CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks

  • 让大模型先用思维链规划,再生成同质高质合成数据。
  • 在MATH500等测试集上超越s1k和OpenMathReasoning数据集。
  • 适合需要强推理和精准指令理解的模型训练场景。

我们提出CoT-Self-Instruct,一种合成数据生成方法:让大模型基于给定种子任务,通过思维链(CoT)进行推理与规划,生成质量与复杂度相当的新合成样本,随后利用自动指标过滤高质量数据用于大模型训练。在可验证推理任务中,该方法生成的数据在MATH500、AMC23、AIME24和GPQA-Diamond测试集上的表现显著优于现有训练数据集(如s1k和OpenMathReasoning)。在不可验证的指令遵循任务中,其性能超过人类和标准Self-Instruct生成的数据,在AlpacaEval 2.0和Arena-Hard基准上实现更优表现。

原文摘要 · Abstract (English)

We propose CoT-Self-Instruct, a synthetic data generation method that instructs LLMs to first reason and plan via Chain-of-Thought (CoT) based on given seed tasks, and then generate a new synthetic example of similar quality and complexity. This is followed by a filtering step to select high-quality data using automatic metrics, which are then used for LLM training. In verifiable reasoning, our synthetic data significantly outperforms existing training datasets, such as s1k and OpenMathReasoning, when evaluated on MATH500, AMC23, AIME24, and GPQA-Diamond. For non-verifiable instruction-following tasks, our method surpasses the performance of both human and standard Self-Instruct training data on the AlpacaEval 2.0 and Arena-Hard benchmarks.

合成数据思维链指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。