用9000条合成数据提升大模型跨学科推理能力。
CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning
- 用先进模型生成长链条思维轨迹,解决数据冷启动问题。
- 覆盖8大学科超1000个细粒度主题,扩展领域广度。
- 全自动评估验证,适合研究通用推理与模型训练者。
大语言模型的推理能力依赖高质量推理数据的监督微调和强化学习后训练。但在开放可扩展场景中,面临三大数据挑战:(1)缺乏初始种子数据,难以启动推理策略;(2)现有开源数据集中于数学,科学领域覆盖有限;(3)前沿推理任务标注成本过高。为此,我们提出CHIMERA,一个包含9000条样本的紧凑合成推理数据集。其具备三方面特性:(1)由先进推理模型生成丰富、长周期的链式思维轨迹;(2)覆盖8大科学领域,通过模型生成的层次化分类体系涵盖超过1000个细粒度主题;(3)采用全自动化可扩展评估流程,利用强推理模型交叉验证问题有效性和答案正确性。我们使用CHIMERA对4B参数的Qwen3模型进行后训练,尽管数据量小,该模型在多个高难度推理基准上表现优异,包括GPQA-Diamond、AIME 24/25/26、HMMT 25和Humanity's Last Exam,性能接近甚至媲美更大规模模型如DeepSeek-R1和Qwen3-235B。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have recently exhibited remarkable reasoning capabilities, largely enabled by supervised fine-tuning (SFT)- and reinforcement learning (RL)-based post-training on high-quality reasoning data. However, reproducing and extending these capabilities in open and scalable settings is hindered by three fundamental data-centric challenges: (1) the cold-start problem, arising from the lack of seed datasets with detailed, long Chain-of-Thought (CoT) trajectories needed to initialize reasoning policies; (2) limited domain coverage, as most existing open-source reasoning datasets are concentrated in mathematics, with limited coverage of broader scientific disciplines; and (3) the annotation bottleneck, where the difficulty of frontier-level reasoning tasks makes reliable human annotation prohibitively expensive or infeasible. To address these challenges, we introduce CHIMERA, a compact synthetic reasoning dataset comprising 9K samples for generalizable cross-domain reasoning. CHIMERA is constructed with three key properties: (1) it provides rich, long CoT reasoning trajectories synthesized by state-of-the-art reasoning models; (2) it has broad and structured coverage, spanning 8 major scientific disciplines and over 1K fine-grained topics organized via a model-generated hierarchical taxonomy; and (3) it employs a fully automated, scalable evaluation pipeline that uses strong reasoning models to cross-validate both problem validity and answer correctness. We use CHIMERA to post-train a 4B Qwen3 model. Despite the dataset's modest size, the resulting model achieves strong performance on a suite of challenging reasoning benchmarks, including GPQA-Diamond, AIME 24/25/26, HMMT 25, and Humanity's Last Exam, approaching or matching the reasoning performance of substantially larger models such as DeepSeek-R1 and Qwen3-235B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。