构建了200万条带多维度标注的思维链数据集,提升大模型推理能力。
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
- 用双大模型生成并验证200万条思维链过程
- 引入冗余度与认知难度新指标,优化模型理解效果
- 适合需要强推理能力的模型训练与评估
大型推理模型(LRMs)在数学求解和代码生成等复杂任务中表现优异,依赖思维链(CoT)过程模拟人类推理。然而,现有CoT数据集缺乏足够数量、高质量的推理样本,且未涵盖描述思维链内部特征的多维属性。为此,我们提出OmniThought,一个包含200万条由两个强大LRM作为教师模型生成并验证的CoT过程的数据集。每条思维链均标注了新的推理冗余度(RV)与认知难度(CD)分数,用于衡量其对模型理解的适配性。我们建立了自洽的数据构建流程。基于Qwen2.5系列模型的实验表明,所提评分可显著提升大模型训练效果。在此基础上,我们训练并发布了多个具备更强推理能力、最优思维链长度与难度的高性能LRM。本工作大幅推动了复杂任务下大型推理模型的发展与训练。
原文摘要 · Abstract (English)
The emergence of large reasoning models (LRMs) has transformed Natural Language Processing by excelling in complex tasks such as mathematical problem-solving and code generation. These models leverage chain-of-thought (CoT) processes, enabling them to emulate human-like reasoning strategies. However, the advancement of LRMs is hindered by the lack of comprehensive CoT datasets. Current resources often fail to provide extensive reasoning problems with coherent CoT processes distilled from multiple teacher models and do not account for multifaceted properties describing the internal characteristics of CoTs. To address these challenges, we introduce OmniThought, a large-scale dataset featuring 2 million CoT processes generated and validated by two powerful LRMs as teacher models. Each CoT process in OmniThought is annotated with novel Reasoning Verbosity (RV) and Cognitive Difficulty (CD) scores, which describe the appropriateness of CoT verbosity and cognitive difficulty level for models to comprehend these reasoning processes. We further establish a self-reliant pipeline to curate this dataset. Extensive experiments using Qwen2.5 models of various sizes demonstrate the positive impact of our proposed scores on LRM training effectiveness. Based on the proposed OmniThought dataset, we further train and release a series of high-performing LRMs, specifically equipped with stronger reasoning abilities and optimal CoT output length and difficulty level. Our contributions significantly enhance the development and training of LRMs for solving complex tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。