生成带口吃特征的车载对话数据,提升AI对真实驾驶场景对话的理解能力。
DRIVE: Disfluency-Rich Synthetic Dialog Data Generation Framework for Intelligent Vehicle Environments
- 用两阶段提示驱动框架动态合成含停顿、重复等口吃行为的对话。
- 在多个基准上优于或媲美训练模型,最高提升BLEU-4 0.61、BERTScore F1 3.48。
- 适合用于训练和增强车载对话系统,尤其在数据稀缺时效果显著。
随着自动驾驶和智能助手普及,车载对话AI日益重要。现有数据集未能捕捉真实驾驶者与AI交互中的自发口吃现象,如犹豫、重述、自修正等。为此,我们提出DiscoDrive,一个包含3500条多轮对话的合成语料库,覆盖七个汽车应用场景,通过两阶段提示驱动的合成管道动态注入口吃。实验表明,DiscoDrive作为训练数据可使DialoGPT-Medium和T5-Base在MultiWOZ 2.2与Schema-Guided Dialogue(SGD)测试集上达到或超越基于KVRET训练的模型表现,最高提升BLEU-4 0.61、METEOR +2.10、ROUGE-L +3.48、BERTScore F1 +3.48。在低资源场景下,与10%的KVRET数据结合使用,额外带来BLEU-4 +0.38、METEOR +1.95、ROUGE-L +2.87、BERTScore F1 +4.00的增益。人工评估显示,DiscoDrive生成的对话在自然度(3.8 vs 3.6)和连贯性(4.1 vs 4.0)上优于KVRET的真实对话,且更符合上下文,优于主流后处理方法(如LARD),同时保持清晰度。DiscoDrive填补了当前资源空白,是训练与增强对话AI的通用语料库。
原文摘要 · Abstract (English)
In-car conversational AI is becoming increasingly critical as autonomous vehicles and smart assistants gain widespread adoption. Yet, existing datasets fail to capture the spontaneous disfluencies such as hesitations, false starts, repetitions, and self-corrections that characterize real driver-AI dialogs. To address this, we introduce DiscoDrive, a synthetic corpus of 3500 multi-turn dialogs across seven automotive domains, generated using a two-stage, prompt-driven pipeline that dynamically integrates disfluencies during synthesis. We show that DiscoDrive is effective both as a training resource, enabling DialoGPT-Medium and T5-Base to match or exceed KVRET-trained models on the MultiWOZ 2.2 and Schema-Guided Dialogue (SGD) relevant test sets (BLEU-4 improvements of 0.26 to 0.61; METEOR +2.10; ROUGE-L +3.48; BERTScore F1 improvements of 1.35 to 3.48), and as a data augmentation resource in low-resource scenarios, delivering additional gains of up to BLEU-4 +0.38, METEOR +1.95, ROUGE-L +2.87, and BERTScore F1 +4.00 when combined with 10 percent of KVRET. Human evaluations further confirm that dialogs sampled from DiscoDrive are rated higher than KVRET's human-collected dialogs in naturalness (3.8 vs 3.6) and coherence (4.1 vs 4.0), and are perceived as more context-appropriate than leading post-hoc methods (such as LARD), without compromising clarity. DiscoDrive fills a critical gap in existing resources and serves as a versatile corpus for both training and augmenting conversational AI, enabling robust handling of real-world, disfluent in-car interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。