揭示大模型推理的脆弱性:超出训练数据范围就失效
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- 从数据分布角度分析思维链为何有效或失效
- 在分布外测试时推理能力急剧下降,证明其不具泛化性
- 适合关注模型可靠性与推理边界的研究者
思维链(CoT)提示已被证明能激发大语言模型的结构化推理能力。尽管广泛应用,但近期研究发现其在某些任务中会失败,引发对CoT推理本质的根本疑问。本文提出数据分布视角,假设CoT推理是模型从分布内数据中学到的结构化归纳偏置,使其能够条件生成近似训练中观察到的推理路径。因此,CoT推理的有效性本质上由训练数据与测试查询之间的分布差异决定。基于此视角,我们从任务、长度和格式三个维度剖析CoT推理。为验证假设,我们引入DataAlchemy——一个抽象且完全可控的环境,从零训练模型并系统探测其在不同分布条件下的表现。通过严谨的受控实验,我们发现当推理超出训练分布时,CoT推理成为脆弱的幻象,凸显实现真正可泛化推理的持续挑战。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) prompting has been shown to be effective in eliciting structured reasoning (i.e., CoT reasoning) from large language models (LLMs). Regardless of its popularity, recent studies expose its failures in some reasoning tasks, raising fundamental questions about the nature of CoT reasoning. In this work, we propose a data distribution lens to understand when and why CoT reasoning succeeds or fails. We hypothesize that CoT reasoning reflects a structured inductive bias learned from in-distribution data, enabling models to conditionally generate reasoning trajectories that approximate those observed during training. As such, the effectiveness of CoT reasoning is fundamentally governed by the nature and degree of distribution discrepancy between training data and test queries. Guided by this lens, we dissect CoT reasoning via three dimensions: task, length, and format. To test the hypothesis, we introduce DataAlchemy, an abstract and fully controllable environment that trains LLMs from scratch and systematically probes them under various distribution conditions. Through rigorous controlled experiments, we reveal that CoT reasoning is a brittle mirage when it is pushed beyond training distributions, emphasizing the ongoing challenge of achieving genuine and generalizable reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。