用微调语音活动投影模型实时预测对话中的'嗯''哦'等回应
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
- 基于语音活动投影模型,连续帧级预测回应时机与类型
- 在真实非平衡数据上实现比基线更优的时序与类型准确率
- 适合需要自然互动的虚拟助手、机器人等语音对话系统
在人类对话中,'yeah'和'oh'等简短回应对促进流畅、有吸引力的交流至关重要。它们表明注意力与理解,又不打断说话人,因此准确预测这些回应对打造更自然的对话代理极为关键。本文提出一种新颖方法,利用微调的语音活动投影(VAP)模型实现实时、连续的回应预测。现有方法多依赖于分轮次或人为平衡的数据集,而本方法在未平衡的真实数据上,以连续帧的方式预测回应的时间与类型。首先在通用对话语料上预训练VAP模型以捕捉对话动态,再在专门关注回应行为的数据集上进行微调。实验结果表明,该模型在时序与类型预测任务上均优于基线方法,在实时环境中表现稳健。这项研究为构建更响应迅速、类人的对话系统提供了可行路径,对虚拟助手、机器人等交互式语音对话应用具有重要意义。
原文摘要 · Abstract (English)
In human conversations, short backchannel utterances such as "yeah" and "oh" play a crucial role in facilitating smooth and engaging dialogue. These backchannels signal attentiveness and understanding without interrupting the speaker, making their accurate prediction essential for creating more natural conversational agents. This paper proposes a novel method for real-time, continuous backchannel prediction using a fine-tuned Voice Activity Projection (VAP) model. While existing approaches have relied on turn-based or artificially balanced datasets, our approach predicts both the timing and type of backchannels in a continuous and frame-wise manner on unbalanced, real-world datasets. We first pre-train the VAP model on a general dialogue corpus to capture conversational dynamics and then fine-tune it on a specialized dataset focused on backchannel behavior. Experimental results demonstrate that our model outperforms baseline methods in both timing and type prediction tasks, achieving robust performance in real-time environments. This research offers a promising step toward more responsive and human-like dialogue systems, with implications for interactive spoken dialogue applications such as virtual assistants and robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。