构建首个土耳其语对话轮换数据集,用于预测说话人何时结束发言。
Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction

- 采用多模态数据(视频、音频、转录)捕捉自然对话中的轮换行为。
- 通过遗传算法优化可解释的决策规则,准确率超基线方法。
- 适合研究多模态对话系统或土耳其语语音处理的研究者。
轮换是人类对话的基本组织特征,但在自然同步对话系统中仍难以建模。尽管已有研究探索多模态方法和大语言模型用于发言结束预测,但针对土耳其语对话轮换动态的自然对话语料库仍属空白。本文引入一个包含未脚本化双人互动的多模态土耳其语对话数据集,包含同步正面视频、按说话人分离的音频通道(可区分重叠语音),以及时间对齐的转录文本。将轮换预测建模为二分类问题,并使用遗传算法(GA)优化由视觉、声学和语言特征推导出的可解释决策规则。提出一种混合的与-或规则表示框架,以表达导致发言转换的多种线索组合。
原文摘要 · Abstract (English)
Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While existing research has explored multimodal approaches and large language models for turn-ending prediction, there is a lack of naturalistic conversational corpora specifically addressing turn-taking dynamics in Turkish. This study introduces a multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions. Turn-taking prediction is formulated as a binary classification problem, and a Genetic Algorithm (GA) is employed to optimize interpretable decision rules derived from visual, acoustic, and linguistic features. A hybrid AND-OR rule representation is adopted in the proposed framework to represent the alternative cue combinations that precede a turn transition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。