提出新融合方法,提升自然语音识别在真实场景下的准确率。
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus

- 用深度交叉注意力融合自监督特征,优化语音表示
- 在FSC Phase-4数据集上实现1.1%的词错误率降低
- 为跨学科研究者提供高质量语音资源支持
自监督学习(SSL)模型显著提升了下游语音任务性能,超越传统人工特征。本研究聚焦大规模自然语音数据集Fearless Steps(FS)APOLLO,特别是FS Challenge(FSC)Phase-4语料库,首次对该数据集进行分析。同时引入CHiME-6数据集,评估在多种自然语音场景下的表现。尽管先前提出的特征精炼损失与融合方法在该语料库上效果有限,本文提出一种新型深度交叉注意力(DCA)融合方法,有效提升性能,尤其针对FSC Phase-4语料库。目标是推动构建更优的FS APOLLO社区资源,满足多学科研究需求。所提方案在词错误率(WER)上取得绝对+1.1%的改进,为大规模语料库提供有效元数据支持。
原文摘要 · Abstract (English)
Using self-supervised learning (SSL) models has significantly improved performance for downstream speech tasks, surpassing the capabilities of traditional hand-crafted features. This study investigates the amalgamation of SSL models, with the aim to leverage both their individual strengths and refine extracted features to achieve improved speech recognition models for naturalistic scenarios. Our research investigates the massive naturalistic Fearless Steps (FS) APOLLO resource, with particular focus on the FS Challenge (FSC) Phase-4 corpus, providing the inaugural analysis of this dataset. Additionally, we incorporate the CHiME-6 dataset to evaluate performance across diverse naturalistic speech scenarios. While exploring previously proposed Feature Refinement Loss and fusion methods, we found these methods to be less effective on the FSC Phase-4 corpus. To address this, we introduce a novel deep cross-attention (DCA) fusion method, designed to elevate performance, especially for the FSC Phase-4 corpus. Our objective is to foster creation of superior FS APOLLO community resources, catering to the diverse needs of researchers across various disciplines. The proposed solution achieves an absolute +1.1% improvement in WER, providing effective meta-data creation for the massive FS APOLLO community resource.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。