arXiv:2411.01483cs.CL2024-11AAAI被引 1

让大模型通过对话反馈学习约束生成,提升可控性。

Teaching Models to Improve on Tape

  • 用强化学习模拟交互,根据约束满足度奖励模型。
  • 在多个任务上优于无反馈的基线方法,提升显著。
  • 支持元学习,可泛化到新任务,适合需精准控制场景。

大语言模型在生成受特定约束内容时常表现不佳,但此类约束通常易于验证是否满足。近期研究显示,模型可从“纠正性反馈”中受益。本文主张,通过训练可显著增强模型利用此类反馈的能力。提出一种基于强化学习的框架CORGI(带引导交互的受控生成),通过模拟交互会话,依据模型满足约束的能力给予奖励。在无需标注数据的情况下,评估多种受控生成任务。结果表明,CORGI持续优于不引入对话反馈的基线强化学习方法。此外,其交互式框架支持元学习,使模型在新任务中的引导交互能力更强。结果明确显示,结合强化学习的对话优化能显著提升大模型在受控生成中的有效性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often struggle when prompted to generate content under specific constraints. However, in such cases it is often easy to check whether these constraints are satisfied or violated. Recent works have shown that LLMs can benefit from such "corrective feedback". Here we claim that this skill of LLMs can be significantly enhanced via training. We introduce an RL framework for teaching models to use such rewards, by simulating interaction sessions, and rewarding the model according to its ability to satisfy the constraints. We refer to our method as CORGI (Controlled Generation with RL for Guided Interaction), and evaluate it on a variety of controlled generation tasks using unlabeled training data. We find that CORGI consistently outperforms the baseline reinforcement learning method that does not incorporate conversational feedback. Furthermore, CORGI's interactive framework enables meta-learning, allowing the LLM to generalize better to guided interaction in new tasks. Our results clearly show that conversational optimization, when combined with reinforcement learning, significantly improves the effectiveness of LLMs in controlled generation contexts.

大模型强化学习可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。