用可控方式生成需逻辑与数值联合推理的题目,提升大模型综合推理能力
LogiNumSynth: Synthesizing Joint Logical-Numerical Reasoning Problems for Language Models
- 通过自然语言合成器灵活控制任务复杂度,支持逻辑与数值推理协同生成
- 实验发现多模型在联合推理上仍有明显短板,证明该数据集诊断价值
- 适合研究大模型推理能力或需要针对性训练数据的学者使用
联合逻辑-数值推理仍是大语言模型的重大挑战,而现有数据集依赖固定规则集,难以调控任务复杂度,限制了其在评估与训练中的泛化能力。我们提出LogiNumSynth,一个灵活的自然语言问题合成器,可生成需同时具备逻辑推理(如规则推理)与数值推理(如算术计算)能力的任务。该方法支持对推理世界丰富度、逻辑深度及数值计算复杂度的细粒度控制,实现跨难度级别的灵活数据合成。我们展示了三项关键贡献:(1) 合成器——生成完全可控的自然语言联合推理任务;(2) 评估与过程分析——同时评估过程准确率与答案准确率;(3) 针对性训练——利用合成数据提升大模型推理表现。多模型实验揭示了持续存在的逻辑-数值推理弱点,表明LogiNumSynth既可作为诊断工具,也可作为提升综合推理能力的靶向监督数据源。
原文摘要 · Abstract (English)
Joint logical-numerical reasoning remains a major challenge for language models, yet existing datasets rely on fixed rule sets and offer limited control over task complexity, constraining their generalizability for evaluation and training. We present LogiNumSynth, a flexible natural language problem synthesizer that synthesizes tasks requiring proficiency in joint logical reasoning (e.g., rule-based reasoning) and numerical reasoning (e.g., arithmetic computation). LogiNumSynth supports fine-grained control over reasoning world richness, logical reasoning depth, and the complexity of numerical computations, enabling flexible data synthesis across difficulty levels. We demonstrate three key contributions: (1) Synthesizer -- synthesizing fully controllable joint reasoning tasks over natural language; (2) Evaluation & Process Analysis -- evaluating both process accuracy and answer accuracy; (3) Targeted Training -- using synthesized data to enhance LLMs' reasoning performance. Experiments with multiple LLMs highlight persistent weaknesses in logical-numerical reasoning, showing that LogiNumSynth can serve as both a diagnostic tool and a source of targeted supervision for advancing integrated reasoning skills.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。