通过测试时扩展控制难题生成难度,提升模型推理能力。
CoDiQ: Test-Time Scaling for Controllable Difficult Question Generation
- 利用测试时扩展调整推理预算,精细控制问题难度。
- 生成4.4万道竞赛级题目,人类解答率超82%且更难于现有数据集。
- 适合需要高质量难题训练的推理模型研究者使用。
大型推理模型(LRMs)在具有挑战性的竞赛级问题上训练可显著提升性能。然而,现有自动化题目生成方法难以精确控制难度,计算成本高,且无法大规模生成竞赛级题目。本文提出CoDiQ(可控难度问题生成)框架,通过测试时扩展实现细粒度难度控制,同时保证问题可解性。首先,我们发现测试时扩展中推理令牌预算增加会提升难度但降低可解性,并确定了模型生成有效高难度问题的固有上限。随后,基于Qwen3-8B开发CoDiQ-Generator,突破了该上限,特别适用于构造难题。在此基础上构建了包含4.4万条竞赛级问题序列的CoDiQ-Corpus。人工评估显示,这些题目比LiveCodeBench/AIME更具挑战性,且解答率超过82%。在CoDiQ-Corpus上训练的LRMs推理性能显著提升,验证了可控难度问题对推理能力的增强作用。论文开源了CoDiQ-Corpus、CoDiQ-Generator及实现代码,以支持相关研究。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) benefit substantially from training on challenging competition-level questions. However, existing automated question synthesis methods lack precise difficulty control, incur high computational costs, and struggle to generate competition-level questions at scale. In this paper, we propose CoDiQ (Controllable Difficult Question Generation), a novel framework enabling fine-grained difficulty control via test-time scaling while ensuring question solvability. Specifically, first, we identify a test-time scaling tendency (extended reasoning token budget boosts difficulty but reduces solvability) and the intrinsic properties defining the upper bound of a model's ability to generate valid, high-difficulty questions. Then, we develop CoDiQ-Generator from Qwen3-8B, which improves the upper bound of difficult question generation, making it particularly well-suited for challenging question construction. Building on the CoDiQ framework, we build CoDiQ-Corpus (44K competition-grade question sequences). Human evaluations show these questions are significantly more challenging than LiveCodeBench/AIME with over 82% solvability. Training LRMs on CoDiQ-Corpus substantially improves reasoning performance, verifying that scaling controlled-difficulty training questions enhances reasoning capabilities. We open-source CoDiQ-Corpus, CoDiQ-Generator, and implementations to support related research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。