用主动学习减少标注量,自动修正语音系统误判意图
IDALC: A Semi-Supervised Framework for Intent Detection and Active Learning based Correction
- 结合半监督学习与主动学习,自动筛选需标注的错误语句
- 在多个数据集上准确率提升5-10%,宏平均F1提高4-8%
- 仅需标注6-10%的未标注数据,显著降低人工成本
语音控制对话系统因能响应多样用户请求而广受欢迎,其通过预定义的技能或意图完成任务。然而系统存在局限:对已知意图若模型置信度低,则拒绝该语句,需人工标注;随着时间推移,还需从被拒语句中提取新意图以扩展功能。持续标注所有新增意图和拒接语句不现实,亟需降低标注成本。本文提出IDALC(基于主动学习的意图检测与修正框架),一种半监督方法,用于识别用户意图并修正系统拒接的语句,同时最小化人工标注需求。在多个基准数据集上的实证结果表明,该系统优于基线方法,准确率提升5-10%,宏平均F1提升4-8%。令人瞩目的是,整体标注成本仅为可用未标注数据的6-10%。
原文摘要 · Abstract (English)
Voice-controlled dialog systems have become immensely popular due to their ability to perform a wide range of actions in response to diverse user queries. These agents possess a predefined set of skills or intents to fulfill specific user tasks. But every system has its own limitations. There are instances where, even for known intents, if any model exhibits low confidence, it results in rejection of utterances that necessitate manual annotation. Additionally, as time progresses, there may be a need to retrain these agents with new intents from the system-rejected queries to carry out additional tasks. Labeling all these emerging intents and rejected utterances over time is impractical, thus calling for an efficient mechanism to reduce annotation costs. In this paper, we introduce IDALC (Intent Detection and Active Learning based Correction), a semi-supervised framework designed to detect user intents and rectify system-rejected utterances while minimizing the need for human annotation. Empirical findings on various benchmark datasets demonstrate that our system surpasses baseline methods, achieving a 5-10% higher accuracy and a 4-8% improvement in macro-F1. Remarkably, we maintain the overall annotation cost at just 6-10% of the unlabelled data available to the system. The overall framework of IDALC is shown in Fig. 1
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。