用合成数据提升土耳其语对话中的轮次预测,让聊天机器人不抢话。
Syn-TurnTurk: A Synthetic Dataset for Turn-Taking Prediction in Turkish Dialogues

- 用多个Qwen大模型生成模拟真实对话的土耳其语数据集。
- 先进模型在该数据集上达到0.839准确率和0.910 AUC。
- 适合研究多语言对话系统或语音交互的开发者参考。
管理自然对话节奏是基于语音的聊天机器人面临的重要挑战。当前多数系统依赖简单的静音检测,但因人类说话存在不规则停顿,常导致机器人打断用户,破坏对话流畅性。这一问题在土耳其语等缺乏高质量轮次预测数据集的语言中尤为严重。本文提出Syn-TurnTurk,一个利用多个Qwen大语言模型生成的合成土耳其语对话数据集,真实模拟了重叠发言和策略性沉默等口语特征。我们使用多种传统与深度学习模型对该数据集进行了评估,结果显示,先进模型如BI-LSTM和集成模型(LR+RF)在该数据集上取得了0.839的准确率和0.910的AUC分数。这些结果表明,该合成数据集有助于模型理解语言线索,从而实现更自然的土耳其语人机交互。
原文摘要 · Abstract (English)
Managing natural dialogue timing is a significant challenge for voice-based chatbots. Most current systems usually rely on simple silence detection, which often fails because human speech patterns involve irregular pauses. This causes bots to interrupt users, breaking the conversational flow. This problem is even more severe for languages like Turkish, which lack high-quality datasets for turn-taking prediction. This paper introduces Syn-TurnTurk, a synthetic Turkish dialogue dataset generated using various Qwen Large Language Models (LLMs) to mirror real-life verbal exchanges, including overlaps and strategic silences. We evaluated the dataset using several traditional and deep learning architectures. The results show that advanced models, particularly BI-LSTM and Ensemble (LR+RF) methods, achieve high accuracy (0.839) and AUC scores (0.910). These findings demonstrate that our synthetic dataset can have a positive affect for models understand linguistic cues, allowing for more natural human-machine interaction in Turkish.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。