用模仿学习检测对话模型的潜在缺陷行为。
Limitation Learning: Catching Adverse Dialog with GAIL
- 通过专家对话示范训练对话策略与判别器
- 判别器能识别出模型生成对话的异常模式
- 适用于发现各类对话模型的隐藏缺陷
模仿学习是一种在无奖励信号时构建策略的有效方法,可通过专家示范实现。本文将其应用于对话系统,训练出一个根据输入提示生成回复的对话策略,以及一个可区分专家对话与合成对话的判别器。虽然策略表现良好,但判别器揭示了当前对话模型存在的局限性。我们提出该方法可用于识别对话任务中任意数据模型的不良行为,具有普适性。
原文摘要 · Abstract (English)
Imitation learning is a proven method for creating a policy in the absence of rewards, by leveraging expert demonstrations. In this work, we apply imitation learning to conversation. In doing so, we recover a policy capable of talking to a user given a prompt (input state), and a discriminator capable of classifying between expert and synthetic conversation. While our policy is effective, we recover results from our discriminator that indicate the limitations of dialog models. We argue that this technique can be used to identify adverse behavior of arbitrary data models common for dialog oriented tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。