arXiv:2509.19695cs.CLcs.AI2025-09ACL

动态切换快慢思维,让对话系统更聪明地探索

DyBBT: Dynamic Balance via Bandit-inspired Targeting for Dialog Policy with Cognitive Dual-Systems

  • 用双系统模型模拟人类思维:快思考与慢思考
  • 在多领域测试中成功率、效率均达顶尖水平
  • 适合研究智能对话决策与认知启发式方法的人

任务导向型对话系统常依赖静态探索策略,无法适应动态对话上下文,导致探索效率低且表现不佳。本文提出 DyBBT 框架,通过结构化认知状态空间建模对话进展、用户不确定性与槽位依赖关系。设计基于强化学习的元控制器,根据实时认知状态和访问频次,在快速直觉推理(系统1)与缓慢深思熟虑(系统2)间动态切换。在单域与多域基准上的大量实验表明,DyBBT 在成功率、效率和泛化能力上均达到当前最优,人类评估也验证其决策与专家判断高度一致。

原文摘要 · Abstract (English)

Task oriented dialog systems often rely on static exploration strategies that do not adapt to dynamic dialog contexts, leading to inefficient exploration and suboptimal performance. We propose DyBBT, a novel dialog policy learning framework that formalizes the exploration challenge through a structured cognitive state space capturing dialog progression, user uncertainty, and slot dependency. DyBBT proposes a bandit inspired meta-controller that dynamically switches between a fast intuitive inference (System 1) and a slow deliberative reasoner (System 2) based on real-time cognitive states and visitation counts. Extensive experiments on single- and multi-domain benchmarks show that DyBBT achieves state-of-the-art performance in success rate, efficiency, and generalization, with human evaluations confirming its decisions are well aligned with expert judgment.

对话系统双系统探索策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。