arXiv:2504.01211cs.AIcs.GT2025-04被引 2

解决隐藏因素干扰下的说服策略评估难题

Off-Policy Evaluation for Sequential Persuasion Process with Unobserved Confounding

  • 将说服过程建模为部分可观测马尔可夫决策过程,处理未知混淆变量
  • 提出基于近端学习的离线评估方法,仅用观测数据即可评估新策略
  • 适用于真实场景中无法实验时的策略效果预判,适合行为科学与智能系统研究者

本文将贝叶斯说服框架拓展至存在未观测混淆变量的发送者-接收者互动场景。传统模型假设信念更新遵循贝叶斯原则,但现实中隐藏变量常影响接收者的信念形成与决策。我们将此建模为多轮序列决策问题:每轮发送者沟通,接收者也与环境交互,其信念更新受未观测混淆变量影响。通过将其重构为部分可观测马尔可夫决策过程(POMDP),捕捉发送者对信念动态和混淆变量的不完全信息。我们证明,在该POMDP中寻找最优观测策略等价于求解原始说服框架中的最优信号策略。进一步,该重构使近端学习可用于说服过程的离线评估,使发送者仅凭行为策略的观测数据即可评估替代信号策略,无需昂贵的新实验。

原文摘要 · Abstract (English)

In this paper, we expand the Bayesian persuasion framework to account for unobserved confounding variables in sender-receiver interactions. While traditional models assume that belief updates follow Bayesian principles, real-world scenarios often involve hidden variables that impact the receiver's belief formation and decision-making. We conceptualize this as a sequential decision-making problem, where the sender and receiver interact over multiple rounds. In each round, the sender communicates with the receiver, who also interacts with the environment. Crucially, the receiver's belief update is affected by an unobserved confounding variable. By reformulating this scenario as a Partially Observable Markov Decision Process (POMDP), we capture the sender's incomplete information regarding both the dynamics of the receiver's beliefs and the unobserved confounder. We prove that finding an optimal observation-based policy in this POMDP is equivalent to solving for an optimal signaling strategy in the original persuasion framework. Furthermore, we demonstrate how this reformulation facilitates the application of proximal learning for off-policy evaluation in the persuasion process. This advancement enables the sender to evaluate alternative signaling strategies using only observational data from a behavioral policy, thus eliminating the necessity for costly new experiments.

因果推断离线评估贝叶斯说服强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。