微调会削弱大模型的思维链推理能力,导致思考过程更不可信。
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
- 对比不同微调方法对思维链推理的影响
- 平均在四个数据集上,思维链忠实度下降
- 适合关注模型可靠性与安全性的研究者
大型语言模型已成为通用智能的强大工具,展现出先进的自然语言处理能力,广泛应用于多个领域。尽管表现优异,近期研究显示通过强化学习人类反馈(RLHF)、监督微调(SFT)和量化低秩适配器(Q-LoRA)等策略可显著提升模型任务性能。然而,以往工作表明,微调虽带来性能提升,也引发灾难性遗忘、隐私与安全风险等问题。目前尚缺乏对微调如何影响大模型推理能力的系统研究。本文探究了微调对大模型推理能力的影响,重点关注任务特定微调对整体推理能力的作用、对思维链(CoT)推理表现的影响,以及对CoT推理忠实度的后果。研究发现,在四个数据集上,微调使CoT推理的忠实度平均下降,揭示微调可能改变了模型内部机制。
原文摘要 · Abstract (English)
Large language models have emerged as powerful tools for general intelligence, showcasing advanced natural language processing capabilities that find applications across diverse domains. Despite their impressive performance, recent studies have highlighted the potential for significant enhancements in LLMs' task-specific performance through fine-tuning strategies like Reinforcement Learning with Human Feedback (RLHF), supervised fine-tuning (SFT), and Quantized Low-Rank Adapters (Q-LoRA) method. However, previous works have shown that while fine-tuning offers significant performance gains, it also leads to challenges such as catastrophic forgetting and privacy and safety risks. To this end, there has been little to no work in \textit{understanding the impact of fine-tuning on the reasoning capabilities of LLMs}. Our research investigates the effect of fine-tuning on the reasoning abilities of LLMs, addressing critical questions regarding the impact of task-specific fine-tuning on overall reasoning capabilities, the influence of fine-tuning on Chain-of-Thought (CoT) reasoning performance, and the implications for the faithfulness of CoT reasonings. By exploring these dimensions, our study shows the impact of fine-tuning on LLM reasoning capabilities, where the faithfulness of CoT reasoning, on average across four datasets, decreases, highlighting potential shifts in internal mechanisms of the LLMs resulting from fine-tuning processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。