将信念估计转化为生成约束,提升对话策略的可靠性。
BEDA: Belief Estimation as Probabilistic Constraints for Performing Strategic Dialogue Acts
- 用概率约束定义对抗与协作两类对话行为
- 在三个任务中均显著提升成功率,最高增20.6点
- 适合需要精准策略决策的对话系统研究者
策略性对话要求智能体执行特定对话行为,信念估计至关重要。现有方法虽能准确估计信念,但缺乏将其用于生成的系统机制。本文提出BEDA框架,首次形式化对抗与对齐两类核心行为,并通过概率约束约束生成内容。该框架包含世界设定、信念估计算器和条件生成器,分别负责信念推断与符合信念的语句生成。在三个场景下——条件守卫劫匪(CKBG,对抗)、互为好友(MF,合作)和CaSiNo(谈判)——BEDA持续优于强基线:在CKBG任务中,各类骨干模型成功率至少提升5.0点,使用GPT-4.1-nano时提升达20.6点;在MF任务中平均提升9.3点;在CaSiNo任务中达成最优协议。结果表明,将信念估计视为约束,是一种简单且通用的可靠策略对话实现方式。
原文摘要 · Abstract (English)
Strategic dialogue requires agents to execute distinct dialogue acts, for which belief estimation is essential. While prior work often estimates beliefs accurately, it lacks a principled mechanism to use those beliefs during generation. We bridge this gap by first formalizing two core acts Adversarial and Alignment, and by operationalizing them via probabilistic constraints on what an agent may generate. We instantiate this idea in BEDA, a framework that consists of the world set, the belief estimator for belief estimation, and the conditional generator that selects acts and realizes utterances consistent with the inferred beliefs. Across three settings, Conditional Keeper Burglar (CKBG, adversarial), Mutual Friends (MF, cooperative), and CaSiNo (negotiation), BEDA consistently outperforms strong baselines: on CKBG it improves success rate by at least 5.0 points across backbones and by 20.6 points with GPT-4.1-nano; on Mutual Friends it achieves an average improvement of 9.3 points; and on CaSiNo it achieves the optimal deal relative to all baselines. These results indicate that casting belief estimation as constraints provides a simple, general mechanism for reliable strategic dialogue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。