arXiv:2503.16544cs.CLcs.AI2025-03被引 5

用因果发现和反事实推理优化对话说服策略,提升系统适应力。

Causal Discovery and Counterfactual Reasoning to Optimize Persuasive Dialogue Policies

  • 通过因果发现识别用户策略对系统回应的影响
  • 构建反事实话语生成对抗网络,提升策略选择效果
  • 适合研究对话系统优化与强化学习的学者参考

个性化说服对话能提升说服效果,但现有系统难以适应用户状态的动态变化。本文提出一种结合因果发现与反事实推理的新方法,以优化系统说服能力。采用GRaSP算法识别用户话语策略(作为状态)与系统策略(作为动作)间的因果关系,将用户策略视为影响系统回应的因果因素,并以此指导双向条件生成对抗网络(BiCoGAN)生成系统反事实话语。随后,利用双延迟深度Q网络(D3QN)模型基于反事实数据选择最优系统话语策略。在PersuasionForGood数据集上的实验表明,该方法相比基线显著提升说服效果:累积奖励与Q值均明显增长,验证了因果发现对增强反事实推理和优化强化学习策略的有效性。

原文摘要 · Abstract (English)

Tailoring persuasive conversations to users leads to more effective persuasion. However, existing dialogue systems often struggle to adapt to dynamically evolving user states. This paper presents a novel method that leverages causal discovery and counterfactual reasoning for optimizing system persuasion capability and outcomes. We employ the Greedy Relaxation of the Sparsest Permutation (GRaSP) algorithm to identify causal relationships between user and system utterance strategies, treating user strategies as states and system strategies as actions. GRaSP identifies user strategies as causal factors influencing system responses, which inform Bidirectional Conditional Generative Adversarial Networks (BiCoGAN) in generating counterfactual utterances for the system. Subsequently, we use the Dueling Double Deep Q-Network (D3QN) model to utilize counterfactual data to determine the best policy for selecting system utterances. Our experiments with the PersuasionForGood dataset show measurable improvements in persuasion outcomes using our approach over baseline methods. The observed increase in cumulative rewards and Q-values highlights the effectiveness of causal discovery in enhancing counterfactual reasoning and optimizing reinforcement learning policies for online dialogue systems.

对话系统因果推理强化学习反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。