提出SpatialEmb模块,直接在端到端ASR中编码空间信息
SpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays
- 设计轻量嵌入模块SpatialEmb,直接提取并编码空间特征
- 在105小时数据上达到17.04%和20.32%字符错误率,刷新纪录
- 无需依赖麦克风布局或说话人位置,适配任意阵列设备
空间信息是多通道多说话人目标语音识别的关键线索。现有先进系统仅在语音分离阶段提取空间特征,随后对分离后的语音进行标准单通道ASR,导致流程冗长且性能受限于预处理模块的误差累积。此外,多数空间特征提取方法依赖说话人位置与麦克风拓扑知识,难以适应新设备。本文提出轻量级嵌入模块SpatialEmb,直接为ASR模型提取并编码空间信息,支持固定与任意麦克风拓扑。我们在真实会议语料AliMeeting上开展全面实验,确定SpatialEmb的最佳模型设计。最佳模型在105小时训练数据(Train-Ali-far)下,评估集与测试集的字符错误率(CER)分别为17.04%和20.32%,在相同训练数据条件下创下新纪录。
原文摘要 · Abstract (English)
Spatial information is a critical clue for multi-channel multi-speaker target speech recognition. Most state-of-the-art multi-channel Automatic Speech Recognition (ASR) systems extract spatial features only during the speech separation stage, followed by standard single-channel ASR on the separated speech. This approach results in an inefficient, lengthy pipeline and sub-optimal ASR performance due to the accumulated errors from preprocessing modules. Furthermore, most spatial feature extraction methods depend on the knowledge of speaker positions and microphone topology, making the systems reliant on specific settings and challenging to adapt to new equipment. In this work, we propose a solution to these issues with a lightweight embedding module named SpatialEmb, which extracts and encodes spatial information directly for the ASR model, supporting both fixed and arbitrary microphone topology. We conduct comprehensive experiments on AliMeeting, a real meeting corpus, to determine the optimal model design for SpatialEmb in terms of both performance and efficiency. Our best model trained with 105 hours Train-Ali-far achieves 17.04% and 20.32% character error rates (CER) on the Eval and Test sets, establishing a new state-of-the-art result with the same training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。