arXiv:2505.06987cs.CLcs.AI2025-05ACL被引 5

用强化学习让大模型学会长远规划,提升情感支持对话质量

Convert Language Model into a Value-based Strategic Planner

  • 将Q-learning融入大模型,实现对话策略的长期优化
  • 在多个数据集上优于直接推理、思维链等主流方法
  • 适合需要长期情感关怀的对话系统研发者

情感支持对话(ESC)旨在通过有效对话缓解个体情绪困扰。尽管大语言模型(LLMs)在该领域取得显著进展,但多数研究未从状态模型角度定义框架,导致长期满意度不足。为解决此问题,我们引入Q-learning于LLMs,提出名为straQ*的框架。该框架使可插拔的LLM能在ESC中自主启动规划,基于长期回报确定最优策略,并最终指导模型生成回应。大量实验表明,straQ*在多个ESC数据集上优于多种基线方法,包括直接推理、自修正、思维链、微调及有限状态机。

原文摘要 · Abstract (English)

Emotional support conversation (ESC) aims to alleviate the emotional distress of individuals through effective conversations. Although large language models (LLMs) have obtained remarkable progress on ESC, most of these studies might not define the diagram from the state model perspective, therefore providing a suboptimal solution for long-term satisfaction. To address such an issue, we leverage the Q-learning on LLMs, and propose a framework called straQ*. Our framework allows a plug-and-play LLM to bootstrap the planning during ESC, determine the optimal strategy based on long-term returns, and finally guide the LLM to response. Substantial experiments on ESC datasets suggest that straQ* outperforms many baselines, including direct inference, self-refine, chain of thought, finetuning, and finite state machines.

情感对话强化学习大模型规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。