研究推理模型在不同任务拓扑下的泛化能力,发现其在宽泛任务中表现好,深度任务中则不如基线。
Boule or Baguette? A Study on Task Topology, Length Generalization, and the Benefit of Reasoning Traces
- 提出任务深度与广度概念,用大规模逻辑命题数据集测试推理痕迹模型的泛化性。
- 在宽而浅的任务上推理痕迹模型表现优于基线,在窄而深的任务上反而更差。
- 揭示推理痕迹机制的本质优势与局限,适用于理解模型在复杂逻辑推理中的表现。
近年来,生成中间推理痕迹(RTs)的神经网络推理模型发展迅速。然而,我们对推理痕迹如何支持推理,以及该范式的局限性仍缺乏完整理解。为此,我们引入PITA:一个包含超过2300万条命题逻辑语句及其对应证明的大规模数据集。作为鲁棒推理的基准,我们关注长度泛化问题:若模型在训练时仅处理最多固定长度的证明,其对更长证明的任务泛化能力如何?我们提出任务深度(解决例题所需步骤数)和任务广度(任务中唯一样本数)两个概念,并在PITA的不同子集上进行测试。结果表明,推理痕迹模型在广度大、深度浅的任务上表现良好,但在广度小、深度大的任务上反而劣于非推理痕迹基线。通过与基于三段论的合成任务对比,我们得出理论结论:推理痕迹模型在深度任务上存在根本性性能上限,而在广度任务上具有显著泛化优势。整体发现揭示了使用推理痕迹的根本性优劣。
原文摘要 · Abstract (English)
Recent years have witnessed meteoric progress in reasoning models: neural networks that generate intermediate reasoning traces (RTs) before producing a final output. Despite the rapid advancement, our understanding of how RTs support reasoning, and the limits of this paradigm, remain incomplete. To promote greater clarity, we introduce PITA: a novel large-scale dataset of over 23 million statements in propositional logic and their corresponding proofs. As a benchmark for robust reasoning, we focus on length generalization: if a model is trained to determine truth or falsity on statements with proofs up to fixed length, how well does it generalize to statements requiring longer proofs? We propose notions of (1) task depth and (2) task breadth, which measure respectively (1) the number of steps required to solve an example from a task and (2) the number of unique examples across a task. We vary these quantities across subsets of PITA, and find that RT models generalize well on broad and shallow subsets, while deteriorating on narrow and deep subsets relative to non-RT baselines. To determine whether our results are idiosyncratic to PITA or indicative of general phenomena, we compare our results to a simple synthetic task based on syllogisms. Our resulting theory suggests fundamental scalings that limit how well RT models perform on deep tasks, and highlights their generalization strengths on broad tasks. Our findings overall identify fundamental benefits and limitations inherent in using reasoning traces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。