用预训练音视频编码器提升语音交互中的发言权预测能力
Voice Activity Projection Model with Multimodal Encoders
- 融合预训练音频与人脸编码器,捕捉细微非语言信号
- 在多个对话指标上达到或超越当前最优模型表现
- 适合研究多模态人机交互与语音行为建模的开发者
对话轮换管理对社交互动至关重要,但人机交互因社会语境复杂且具多模态特性而难以建模。传统系统依赖静默时长,而现有语音活动投影(VAP)模型通过统一建模对话轮换行为作为预测目标,显著提升了预测性能。近期提出的多模态VAP模型已大幅超越此前最优方案。本文提出一种增强型多模态模型,引入预训练的音频与面部编码器,以捕捉更细微的表达特征。实验表明,该模型在多个轮换预测指标上表现优异,部分场景下甚至超越现有最先进模型。全部源代码与预训练模型已开源至https://github.com/sagatake/VAPwithAudioFaceEncoders。
原文摘要 · Abstract (English)
Turn-taking management is crucial for any social interaction. Still, it is challenging to model human-machine interaction due to the complexity of the social context and its multimodal nature. Unlike conventional systems based on silence duration, previous existing voice activity projection (VAP) models successfully utilized a unified representation of turn-taking behaviors as prediction targets, which improved turn-taking prediction performance. Recently, a multimodal VAP model outperformed the previous state-of-the-art model by a significant margin. In this paper, we propose a multimodal model enhanced with pre-trained audio and face encoders to improve performance by capturing subtle expressions. Our model performed competitively, and in some cases, even better than state-of-the-art models on turn-taking metrics. All the source codes and pretrained models are available at https://github.com/sagatake/VAPwithAudioFaceEncoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。