研究数据分布偏移如何破坏思维链推理效果
Data Shifts Hurt CoT: A Theoretical Study
- 分析思维链在分布偏移与数据污染下的表现退化机制
- 发现思维链反而比直接预测更差,尤其在奇偶性问题上
- 为模型鲁棒性提供理论解释,适合关注AI可靠性研究者
思维链(CoT)已被证明能有效提升大语言模型的输出质量,尤其在解决如k-奇偶性这类计算难题时表现出色。然而现有研究依赖于训练与测试分布一致、数据无污染等理想假设,现实环境中这些条件常不成立。本文首次系统研究数据偏移对最佳实践下CoT方法的损害,聚焦于k-奇偶性问题,考察分布偏移与数据投毒的联合影响。结果揭示一个意外现象:使用思维链反而导致学习奇偶性的性能劣于直接生成预测。技术分析进一步提供了该现象的严谨机理解释,阐明了数据偏移如何破坏思维链的推理路径。
原文摘要 · Abstract (English)
Chain of Thought (CoT) has been applied to various large language models (LLMs) and proven to be effective in improving the quality of outputs. In recent studies, transformers are proven to have absolute upper bounds in terms of expressive power, and consequently, they cannot solve many computationally difficult problems. However, empowered by CoT, transformers are proven to be able to solve some difficult problems effectively, such as the $k$-parity problem. Nevertheless, those works rely on two imperative assumptions: (1) identical training and testing distribution, and (2) corruption-free training data with correct reasoning steps. However, in the real world, these assumptions do not always hold. Although the risks of data shifts have caught attention, our work is the first to rigorously study the exact harm caused by such shifts to the best of our knowledge. Focusing on the $k$-parity problem, in this work we investigate the joint impact of two types of data shifts: the distribution shifts and data poisoning, on the quality of trained models obtained by a well-established CoT decomposition. In addition to revealing a surprising phenomenon that CoT leads to worse performance on learning parity than directly generating the prediction, our technical results also give a rigorous and comprehensive explanation of the mechanistic reasons of such impact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。