arXiv:2410.19920cs.LG2024-10NAACL被引 7

RL训练后大模型对提示敏感,我们提出方法提升泛化能力。

Reinforcement Learning for Aligning Large Language Models Agents with Interactive Environments: Quantifying and Mitigating Prompt Overfitting

  • 通过对比损失降低模型对提示格式的依赖
  • 发现训练时提示变化会导致性能下降
  • 适用于需要稳定交互的智能体应用

强化学习(RL)是将大语言模型(LLM)知识与序列决策任务对齐的有前景方法。然而,极少研究系统考察在特定环境中通过RL微调对LLM智能体能力的影响。本文提出新框架,分析LLM在文本环境中经RL训练后对提示形式的敏感性。结果表明,当面对与训练阶段不同的提示形式时,LLM性能显著下降。进一步通过分析模型内部表示和显著标记,揭示敏感性来源。最后,我们引入对比损失来缓解此问题,提升模型鲁棒性和泛化能力。

原文摘要 · Abstract (English)

Reinforcement learning (RL) is a promising approach for aligning large language models (LLMs) knowledge with sequential decision-making tasks. However, few studies have thoroughly investigated the impact on LLM agents capabilities of fine-tuning them with RL in a specific environment. In this paper, we propose a novel framework to analyze the sensitivity of LLMs to prompt formulations following RL training in a textual environment. Our findings reveal that the performance of LLMs degrades when faced with prompt formulations different from those used during the RL training phase. Besides, we analyze the source of this sensitivity by examining the model's internal representations and salient tokens. Finally, we propose to use a contrastive loss to mitigate this sensitivity and improve the robustness and generalization capabilities of LLMs.

强化学习大模型提示敏感泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。