让对话模型根据用户性格动态调整说服策略,提升说服效果。
Personality-Aware Reinforcement Learning for Persuasive Dialogue with LLM-Driven Simulation
- 基于议程的框架控制策略选择,用MMR保证回复相关性和多样性。
- 每轮生成81维性格嵌入,结合对话历史优化决策,累计奖励提升23%。
- 适合研究个性化对话系统或行为建模的开发者与研究员。
有效的说服性对话代理需根据用户个体差异动态调整策略,并考虑其心理状态和意图随对话演化的特征。本文提出一种人格感知强化学习方法,包含三个核心模块:(1) 策略导向交互框架,作为基于议程的策略控制器,通过最大边际相关性(MMR)检索生成上下文相关、多样且可扩展的数据;(2) 人格感知用户表征学习,从近期对话中预测每轮81维混合类型嵌入并加入强化学习状态;(3) 双重双DQN(D3QN)模型与奖励预测,策略基于对话历史和轮次级人格估计,使用包含同意意图、捐款金额及反悔惩罚的复合奖励进行训练。利用基于议程的大型语言模型(LLM)模拟流程生成多样化交互,从中推断人格特征。在增强模拟对话的PersuasionForGood(P4G)数据集上实验表明:(i) 轮次级人格条件显著提升策略适应性与累计说服奖励;(ii) LLM驱动模拟提升对未见用户行为的泛化能力;(iii) 引入反悔惩罚减少同意后的撤回,同时小幅提升捐款结果。结果表明,结构化交互、动态人格估计与行为驱动奖励协同可生成更有效的说服策略。
原文摘要 · Abstract (English)
Effective persuasive dialogue agents adapt their strategies to individual users, accounting for the evolution of their psychological states and intentions throughout conversations. We present a personality-aware reinforcement learning approach comprising three main modules: (1) a Strategy-Oriented Interaction Framework, which serves as an agenda-based strategy controller that selects strategy-level actions and generate responses via Maximal Marginal Relevance (MMR) retrieval to ensure contextual relevance, diversity, and scalable data generation; (2) Personality-Aware User Representation Learning, which produces an 81-dimensional mixed-type embedding predicted at each turn from recent exchanges and appended to the reinforcement learning state; and (3) a Dueling Double DQN (D3QN) model and Reward Prediction, in which the policy is conditioned on dialogue history and turn-level personality estimates and trained using a composite reward incorporating agreement intent, donation amount, and changeof-mind penalties. We use an agenda-based LLM simulation pipeline to generate diverse interactions, from which personality estimation is inferred from the generated utterances. Experiments on the PersuasionForGood (P4G) dataset augmented with simulated dialogues reveal three main findings: (i) turn-level personality conditioning improves policy adaptability and cumulative persuasion rewards; (ii) LLM-driven simulation enhances generalization to unseen user behaviors; and (iii) incorporating a change-of-mind penalty reduces post-agreement retractions while slightly improving donation outcomes. These results demonstrate that structured interaction, dynamic personality estimation, and behaviorally informed rewards together yield more effective persuasive policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。