通过误导大模型,发现其具备抽象推理能力。
Exploring the Hidden Reasoning Process of Large Language Models by Misleading Them
- 用矛盾规则微调模型,测试其是否能脱离记忆进行推理。
- 模型能在未见过的任务中应用错误规则,表现泛化能力。
- 揭示大模型存在先抽象后推理的内在机制,适合研究推理本质者阅读。
大型语言模型(LLMs)在多种场景下展现出推理能力,但它们是真正进行了任务抽象与基于规则的推理,还是仅依赖记忆?为回答这一问题,我们提出一种新颖的实验方法——误导微调(MisFT),通过构建与正确原则相悖的数学表达式或逻辑公式数据集,使模型学习这些矛盾规则,并评估其在未见测试领域中的泛化能力。一系列实验表明,当前的LLMs能够将矛盾规则应用于解决实际的数学应用题和自然语言推理任务,暗示了LLMs内部存在先抽象再推理的机制。
原文摘要 · Abstract (English)
Large language models (LLMs) have been able to perform various forms of reasoning tasks in a wide range of scenarios, but are they truly engaging in task abstraction and rule-based reasoning beyond mere memorization? To answer this question, we propose a novel experimental approach, Misleading Fine-Tuning (MisFT), to examine whether LLMs perform abstract reasoning by altering their original understanding of fundamental rules. In particular, by constructing datasets with math expressions or logical formulas that contradict correct principles, we fine-tune the model to learn those contradictory rules and assess its generalization ability on unseen test domains. Through a series of experiments, we find that current LLMs are capable of applying contradictory rules to solve practical math word problems and natural language reasoning tasks, implying the presence of an internal mechanism in LLMs that abstracts before reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。