arXiv:2510.08585eess.AScs.AI2025-10

用发音特征提升语音识别,尤其在数据少时效果更明显。

Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion

  • 通过语音反演任务预测发音特征,再用交叉注意力融合到模型中。
  • 在LibriSpeech上优于主流Transformer模型,低资源下提升更显著。
  • 适合研究语音识别中的多模态融合或资源受限场景应用。

先前研究将发音特征作为语音识别的补充表示,但多局限于浅层声学模型。本文在深度学习时代重新审视发音信息,提出一种框架:将发音特征既作为辅助预测任务,又作为伪输入注入识别模型。具体而言,采用语音反演作为辅助任务,预测出的发音特征以查询流形式,通过交叉注意力模块与声学嵌入(键值对)融合。在LibriSpeech上的实验表明,该方法在强基线(Transformer)基础上实现持续提升,尤其在低资源条件下表现更优。结果表明,一旦结合现代架构,曾被忽视的发音特征能带来实质性收益。

原文摘要 · Abstract (English)

Prior works have investigated the use of articulatory features as complementary representations for automatic speech recognition (ASR), but their use was largely confined to shallow acoustic models. In this work, we revisit articulatory information in the era of deep learning and propose a framework that leverages articulatory representations both as an auxiliary task and as a pseudo-input to the recognition model. Specifically, we employ speech inversion as an auxiliary prediction task, and the predicted articulatory features are injected into the model as a query stream in a cross-attention module with acoustic embeddings as keys and values. Experiments on LibriSpeech demonstrate that our approach yields consistent improvements over strong transformer-based baselines, particularly under low-resource conditions. These findings suggest that articulatory features, once sidelined in ASR research, can provide meaningful benefits when reintroduced with modern architectures.

语音识别发音特征交叉注意力低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。