arXiv:2506.03980cs.CL2025-06被引 7

用预训练音视频编码器提升语音交互中的发言权预测能力

Voice Activity Projection Model with Multimodal Encoders

  • 融合预训练音频与人脸编码器,捕捉细微非语言信号
  • 在多个对话指标上达到或超越当前最优模型表现
  • 适合研究多模态人机交互与语音行为建模的开发者

对话轮换管理对社交互动至关重要,但人机交互因社会语境复杂且具多模态特性而难以建模。传统系统依赖静默时长,而现有语音活动投影(VAP)模型通过统一建模对话轮换行为作为预测目标,显著提升了预测性能。近期提出的多模态VAP模型已大幅超越此前最优方案。本文提出一种增强型多模态模型,引入预训练的音频与面部编码器,以捕捉更细微的表达特征。实验表明,该模型在多个轮换预测指标上表现优异,部分场景下甚至超越现有最先进模型。全部源代码与预训练模型已开源至https://github.com/sagatake/VAPwithAudioFaceEncoders。

原文摘要 · Abstract (English)

Turn-taking management is crucial for any social interaction. Still, it is challenging to model human-machine interaction due to the complexity of the social context and its multimodal nature. Unlike conventional systems based on silence duration, previous existing voice activity projection (VAP) models successfully utilized a unified representation of turn-taking behaviors as prediction targets, which improved turn-taking prediction performance. Recently, a multimodal VAP model outperformed the previous state-of-the-art model by a significant margin. In this paper, we propose a multimodal model enhanced with pre-trained audio and face encoders to improve performance by capturing subtle expressions. Our model performed competitively, and in some cases, even better than state-of-the-art models on turn-taking metrics. All the source codes and pretrained models are available at https://github.com/sagatake/VAPwithAudioFaceEncoders.

多模态语音交互编码器对话管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。