首个面向智能眼镜的多通道语音基础模型,用自监督学习提升听觉感知能力。
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
- 基于无几何约束的自监督学习,统一建模多麦克风输入信号
- 仅用8小时标注数据,语音识别效果超2000小时训练的监督模型
- 适用于语音识别、声源定位等真实场景,适合可穿戴设备开发者
智能眼镜等多通道可穿戴设备日益普及,催生了定向语音识别和助听等应用。然而现有方法依赖独立训练的模型,难以利用大量未标注数据。本文提出 M-BEST-RQ,首个面向智能眼镜的多通道语音基础模型,采用无几何依赖的自监督学习(SSL)架构。不同于以往仅在模拟环境评估,我们构建了真实场景下的三项下游任务:对话式自动语音识别(ASR)、球面主动声源定位、佩戴者语音活动检测,数据源自 MMCSG 与 EasyCom 数据集。实验表明,通用的 M-BEST-RQ 编码器在所有任务上均达到或超过监督模型表现。尤其在对话式 ASR 上,仅使用 8 小时标注数据即超越需 2000 小时标注的监督基线,验证了该方法的有效性。
原文摘要 · Abstract (English)
The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tasks use independently trained models, which may not benefit from large amounts of unlabeled data. In this paper, we propose M-BEST-RQ, the first multi-channel speech foundation model for smart glasses, which is designed to leverage large-scale self-supervised learning (SSL) in an array-geometry agnostic approach. While prior work on multi-channel speech SSL only evaluated on simulated settings, we curate a suite of real downstream tasks to evaluate our model, namely (i) conversational automatic speech recognition (ASR), (ii) spherical active source localization, and (iii) glasses wearer voice activity detection, which are sourced from the MMCSG and EasyCom datasets. We show that a general-purpose M-BEST-RQ encoder is able to match or surpass supervised models across all tasks. For the conversational ASR task in particular, using only 8 hours of labeled speech, our model outperforms a supervised ASR baseline that is trained on 2000 hours of labeled data, which demonstrates the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。