让大模型从解题经验中提炼可复用的抽象知识,提升推理能力。
Notes to Self: Can LLMs Benefit from Experiential Abstractions?
- 从解题过程提取自然语言抽象知识,构建可检索库
- 在数学与逻辑推理任务上显著提升模型性能
- 模型自提取抽象效果接近人工提炼,通用性强
人类会将经验提炼为可复用的抽象知识,如策略和警示提醒,并逐步提升问题解决效率。本文研究大语言模型(LLMs)是否也能从经验中获益于此类抽象。基于 MATH 训练集上的解题轨迹,由更强的教师或模型自身提取自然语言形式的抽象知识,构建可检索的知识库。探索两种使用方式:(1) 推理时检索抽象;(2) 使用抽象增强提示进行强化学习训练。实验表明,这些经验抽象显著提升了模型在数学与逻辑推理基准上的表现。自提取的抽象效果接近教师提取结果,且该框架可迁移至其他数据集与模型。结果表明,大模型能像人类一样,自主提取并应用经验抽象。
原文摘要 · Abstract (English)
Humans distill experience into reusable abstractions, e.g., strategies and cautionary reminders, and apply them to gradually solve problems more effectively. We study whether Large Language Models (LLMs) can similarly benefit from such experiential abstractions. From LLMs' solution traces on the MATH training set, a stronger teacher or the LLMs themselves extract natural-language abstractions into a retrievable library. We explore two usage modes: (1) inference-time retrieval and (2) reinforcement learning (RL) with abstraction-augmented training prompts. Experiential abstractions improve LLM performance on mathematical and logical reasoning benchmarks. Self-extracted abstractions match teacher-extracted ones, and our abstraction usage framework can transfer to other datasets and models. These findings suggest LLMs can extract and apply experiential abstractions much as humans leverage distilled experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。