arXiv:2607.01557cs.CLcs.AI2026-07

用强化学习动态选说服策略,提升高风险场景下的劝离成功率。

DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents

论文配图:DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents
图 1 · 摘自论文原文
  • 基于对话上下文实时选择说服策略,使用Q-learning优化决策
  • 在模拟与真人测试中,劝离成功率显著高于零样本LLM和通用RAG方法
  • 适用于高风险场景的个性化说服,如火灾疏散等紧急救援

大型语言模型在高风险说服场景中表现不佳,因个体性格与关切各不相同,需定制化策略而非通用方案。为此,我们以火灾救援场景为高风险说服任务,提出对话策略选择(DiPS)框架,采用Q-learning机制,根据居民最近发言动态选择说服策略。通过训练一个能最大化撤离成功率的评估器,实现每轮对话中的策略优选。我们在模拟环境与真实人类交互中对比多个基线方法,结果表明,DiPS在撤离成功率上优于零样本大模型及通用RAG增强方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns require tailored strategies rather than a one-size-fits-all approach. To address this challenge, we focus on a fire-rescue scenario in which an operator must persuade a resident to evacuate as a high-stakes persuasion domain and propose Dialogue Policy Selection (DiPS), a Q-learning framework to dynamically select persuasion strategies adapted to the evolving conversational context. Specifically, we train a critic, trained to maximize the chance of evacuation success, to select a persuasion policy at each turn based on the resident's recent utterances. We then evaluate DiPS against multiple baselines in both simulated and real human interactions. We find that DiPS achieves higher evacuation success than a zero-shot LLM and generic RAG-augmented approach.

对话系统强化学习说服生成高风险场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。