用LLM和语音活动投影融合预测对话轮换,提升准确性
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
- 融合大语言模型与语音活动投影,双模态协同预测
- 在ICC和CCPE数据集上验证,提升非剧本对话轮换识别率
- 适合语音交互系统开发者,提升人机对话自然度
对话轮换预测旨在判断对话中说话人何时将发言权让给他人。本研究通过多模态集成方法,结合大语言模型(LLMs)与语音活动投影(VAP)模型,提升对真实对话场景中轮换点(TRPs)的识别精度与效率。该方法融合语言理解能力与语音时序特征,在剧本化与非剧本化对话中均表现优异。在In-Conversation Corpus(ICC)与Coached Conversational Preference Elicitation(CCPE)数据集上进行评估,揭示现有模型的优劣,并提出更具鲁棒性的预测框架。
原文摘要 · Abstract (English)
Turn-taking prediction is the task of anticipating when the speaker in a conversation will yield their turn to another speaker to begin speaking. This project expands on existing strategies for turn-taking prediction by employing a multi-modal ensemble approach that integrates large language models (LLMs) and voice activity projection (VAP) models. By combining the linguistic capabilities of LLMs with the temporal precision of VAP models, we aim to improve the accuracy and efficiency of identifying TRPs in both scripted and unscripted conversational scenarios. Our methods are evaluated on the In-Conversation Corpus (ICC) and Coached Conversational Preference Elicitation (CCPE) datasets, highlighting the strengths and limitations of current models while proposing a potentially more robust framework for enhanced prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。