arXiv:2409.09284cs.SDcs.MM2024-09被引 1

提出多模态多视角模型,让语音助手更准分辨是否在对自己说话。

M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection

  • 设计多视角网络,融合文本、语音及对齐视图提升鲁棒性
  • 在ASR错误数据上首次超越人类判断准确率
  • 适合智能语音交互系统研发与人机对话优化场景

为实现更自然的人机语音交互,当前研究聚焦于无需重复唤醒词的全双工模式。这要求在复杂声源环境下,语音助手能准确区分话语是否指向设备。目前主流采用文本与语音联合建模的双编码器结构,但因自动语音识别(ASR)误差导致输入不匹配时,模型常出错。为此,我们提出M$^{3}$V,一种多模态多视角的设备指向性语音检测方法,将问题建模为多视角学习任务,在多模态基础上引入单模态视图与文本-音频对齐视图。实验表明,M$^{3}$V显著优于仅使用单模态或多模态训练的模型,并首次在含ASR错误的数据上超越人类判断性能。

原文摘要 · Abstract (English)

With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on repeated wake-up words. This requires that in scenes with complex sound sources, the voice assistant must classify utterances as device-oriented or non-device-oriented. The dual-encoder structure, which is jointly modeled by text and speech, has become the paradigm of device-directed speech detection. However, in practice, these models often produce incorrect predictions for unaligned input pairs due to the unavoidable errors of automatic speech recognition (ASR).To address this challenge, we propose M$^{3}$V, a multi-modal multi-view approach for device-directed speech detection, which frames we frame the problem as a multi-view learning task that introduces unimodal views and a text-audio alignment view in the network besides the multi-modal. Experimental results show that M$^{3}$V significantly outperforms models trained using only single or multi-modality and surpasses human judgment performance on ASR error data for the first time.

语音识别多模态人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。