让大模型学会根据语言反馈动态调整推理,提升自我改进能力。
Improving Interactive In-Context Learning from Natural Language Feedback
- 将单轮任务转为多轮教学互动,利用信息差驱动反馈学习。
- 小模型经训练后性能接近大10倍的模型,且跨领域泛化能力强。
- 可实现无外部教师的自我纠错,为模型自进化提供新路径。
人类学习依赖于在协作中根据纠正性反馈调整思维过程,而当前大模型训练主要依赖静态语料库,忽视了动态交互反馈的重要性。本文提出一种框架,将这种交互式上下文学习视为可训练的独立技能。通过将单轮可验证任务转化为由信息不对称驱动的多轮教学互动,我们发现现有主流模型在复杂推理任务中难以吸收纠正性反馈。经过该方法训练后,模型显著提升了从语言反馈中交互学习的能力:小模型的多轮表现几乎达到大10倍模型的水平。同时,数学任务上的交互训练展现出强大的分布外泛化能力,可迁移至编程、谜题和迷宫导航等不同领域。定性分析表明,性能提升源于增强的上下文可塑性。最后,我们证明该范式可实现统一的自我改进:通过训练模型预测教师的批评,有效建模反馈环境,使模型即使无教师也能实现内部自修正。
原文摘要 · Abstract (English)
Adapting one's thought process based on corrective feedback is an essential ability in human learning, particularly in collaborative settings. In contrast, the current large language model training paradigm relies heavily on modeling vast, static corpora. While effective for knowledge acquisition, it overlooks the interactive feedback loops essential for models to adapt dynamically to their context. In this work, we propose a framework that treats this interactive in-context learning ability not as an emergent property, but as a distinct, trainable skill. We introduce a scalable method that transforms single-turn verifiable tasks into multi-turn didactic interactions driven by information asymmetry. We first show that current flagship models struggle to integrate corrective feedback on hard reasoning tasks. We then demonstrate that models trained with our approach dramatically improve the ability to interactively learn from language feedback. More specifically, the multi-turn performance of a smaller model nearly reaches that of a model an order of magnitude larger. We also observe robust out-of-distribution generalization: interactive training on math problems transfers to diverse domains like coding, puzzles and maze navigation. Our qualitative analysis suggests that this improvement is due to an enhanced in-context plasticity. Finally, we show that this paradigm offers a unified path to self-improvement. By training the model to predict the teacher's critiques, effectively modeling the feedback environment, we convert this external signal into an internal capability, allowing the model to self-correct even without a teacher.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。