首次将语音活动预测拓展至三人对话,提升多轮对话流畅性。
Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems
- 基于声学数据预测三人对话中各说话人未来发言状态。
- 三人群体对话下的语音活动预测准确率优于基线模型。
- 适用于需要自然交互的多人语音对话系统开发。
话轮转换是口语对话的核心,但以往研究多集中于两人场景。本文首次将语音活动投影(VAP)方法扩展至三人对话场景,旨在仅利用声学数据预测每位说话人未来的发言状态。研究在包含多种话题讨论的日本三人群体对话数据集上训练了多个模型。结果表明,针对三人群体训练的VAP模型在所有情况下均优于基线,但对话类型对预测准确性有显著影响。本研究证实了VAP可用于三人群体对话中的话轮预测,未来工作将把该模型集成进实际口语对话系统。
原文摘要 · Abstract (English)
Turn-taking is a fundamental component of spoken dialogue, however conventional studies mostly involve dyadic settings. This work focuses on applying voice activity projection (VAP) to predict upcoming turn-taking in triadic multi-party scenarios. The goal of VAP models is to predict the future voice activity for each speaker utilizing only acoustic data. This is the first study to extend VAP into triadic conversation. We trained multiple models on a Japanese triadic dataset where participants discussed a variety of topics. We found that the VAP trained on triadic conversation outperformed the baseline for all models but that the type of conversation affected the accuracy. This study establishes that VAP can be used for turn-taking in triadic dialogue scenarios. Future work will incorporate this triadic VAP turn-taking model into spoken dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。