arXiv:2602.16488cs.CLcs.AI2026-02被引 1

让大模型学会主动问问题,从对话中高效学习。

Learning to Learn from Language Feedback with Social Meta-Learning

  • 用模拟教学对话训练模型主动寻求语言反馈
  • 跨领域提升:数学训练后能更好解决编程问题
  • 面对模糊任务更少乱答,更会主动追问关键信息

大型语言模型在对话中常无法有效获取纠正性反馈,且极少主动请求澄清,导致交流显得呆板。为此,我们受人类社会元学习(SML)启发,将SML转化为微调方法,通过模拟教学对话,训练模型主动求助并利用语言反馈解决问题。该方法使模型在单轮无法求解时,通过多轮互动获得突破。跨任务泛化效果显著:在数学问题上训练的模型,能更好处理编程问题;反之亦然。即使仅在完整问题上训练,这些模型在面对信息逐步披露的模糊任务时,也更少提前作答,更愿意主动询问缺失信息。本研究提出了一种可扩展的AI学习范式,使其能真正从语言反馈中成长。

原文摘要 · Abstract (English)

Large language models (LLMs) often struggle to learn from corrective feedback within a conversational context. They are rarely proactive in soliciting this feedback, even when faced with ambiguity, which can make their dialogues feel static, one-sided, and lacking the adaptive qualities of human conversation. To address these limitations, we draw inspiration from social meta-learning (SML) in humans - the process of learning how to learn from others. We formulate SML as a finetuning methodology, training LLMs to solicit and learn from language feedback in simulated pedagogical dialogues, where static tasks are converted into interactive social learning problems. SML effectively teaches models to use conversation to solve problems they are unable to solve in a single turn. This capability generalises across domains; SML on math problems produces models that better use feedback to solve coding problems and vice versa. Furthermore, despite being trained only on fully-specified problems, these models are better able to solve underspecified tasks where critical information is revealed over multiple turns. When faced with this ambiguity, SML-trained models make fewer premature answer attempts and are more likely to ask for the information they need. This work presents a scalable approach to developing AI systems that effectively learn from language feedback.

语言反馈对话学习元学习交互式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。