arXiv:2602.15484eess.AScs.LG2026-02

用瓶颈Transformer提升语音可懂度预测准确率

Bottleneck Transformer-Based Approach for Improved Automatic STOI Score Prediction

  • 采用瓶颈Transformer架构,结合卷积与自注意力机制提取特征
  • 在已见和未见场景下相关性更高,均方误差更低
  • 参数少却性能优于现有自监督模型,适合实际部署

本研究提出一种基于瓶颈Transformer的新型方法,用于预测短时客观可懂度(STOI)评分。传统STOI计算需干净参考语音,限制了实际应用。为此,众多基于深度学习的非侵入式语音评估模型受到关注。尽管已有研究表现良好,但仍存在提升空间。本文采用瓶颈Transformer架构,融合卷积模块学习帧级特征,并通过多头自注意力层聚合信息,使模型聚焦输入关键内容。实验表明,该模型在已见与未见场景下的相关性更高、均方误差更低,优于使用自监督学习和频谱特征作为输入的当前最优模型。

原文摘要 · Abstract (English)

In this study, we have presented a novel approach to predict the Short-Time Objective Intelligibility (STOI) metric using a bottleneck transformer architecture. Traditional methods for calculating STOI typically requires clean reference speech, which limits their applicability in the real world. To address this, numerous deep learning-based nonintrusive speech assessment models have garnered significant interest. Many studies have achieved commendable performance, but there is room for further improvement. We propose the use of bottleneck transformer, incorporating convolution blocks for learning frame-level features and a multi-head self-attention (MHSA) layer to aggregate the information. These components enable the transformer to focus on the key aspects of the input data. Our model has shown higher correlation and lower mean squared error for both seen and unseen scenarios compared to the state-of-the-art model using self-supervised learning (SSL) and spectral features as inputs.

语音评估TransformerSTOI瓶颈结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。