arXiv:2605.11519cs.AIcs.CL2026-05被引 1

提出因果视角的可控用户模拟,解决传统方法导致评估偏差的问题。

Controllable User Simulation

论文配图:Controllable User Simulation
图 1 · 摘自论文原文
  • 将可控用户模拟建模为因果推断问题,识别出训练标签的结构偏差
  • 发现标准方法在策略变化时会导致评估方差几何级爆炸,称为可控性崩溃
  • 提出先验控制、分步动态控制等新方法,实现零样本泛化且保持对话多样性

使用离线数据集评估对话智能体常无法覆盖罕见场景或测试新策略。为此,研究者采用可控用户模拟器进行目标化、反事实评估,通常通过提示或微调大语言模型实现。本文将可控模拟形式化为因果推断问题,连接自然语言评估与离线策略评估方法,发现通过事后轨迹标签进行监督微调的标准做法会产生结构性偏差。这些标签与数据生成行为策略紧密耦合,引入前瞻偏差,破坏因果一致性。此外,我们证明在策略转移下,该缺陷会导致评估指标方差几何级膨胀,这一现象称为可控性崩溃。为恢复因果一致性,我们建立了准确模拟的理论条件,并提出实用训练缓解策略:先验控制、分步动态控制与直接策略条件学习。实验表明,尽管标准全局控制会扭曲对话分布并丧失行为多样性,我们基于因果框架的模拟器能消除前瞻偏差,保持自然方差,并对未见智能体行为展现出鲁棒的零样本泛化能力。

原文摘要 · Abstract (English)

Using offline datasets to evaluate conversational agents often fails to cover rare scenarios or to support testing new policies. This has motivated the use of controllable user simulators for targeted, counterfactual evaluation, typically implemented by prompting or fine-tuning large language models. In this work, we formalize controllable simulation as a causal inference problem. By bridging natural language evaluation with off-policy evaluation methodology, we show that the standard practice of training simulators via supervised fine-tuning on post-hoc trajectory labels yields a structurally biased model. Specifically, these labels are inextricably coupled to the data-generating behavior policy, injecting a look-ahead bias that breaks causal consistency. Furthermore, we prove that under policy shift this failure causes the variance of evaluation metrics to explode geometrically, a phenomenon we term controllability collapse. To restore causal consistency, we establish theoretical conditions for accurate simulation and propose practical training mitigations: a priori controls, step-wise dynamic controls, and direct policy-conditioned learning. Empirical evaluation confirms that while standard global controls distort conversational distributions and collapse behavioral diversity, our causally grounded simulators eliminate look-ahead bias, preserve natural variance, and exhibit robust zero-shot generalization to unseen agent behaviors.

对话系统因果推断用户模拟评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。