用翻译语义和语音特征增强低资源语音识别,提升准确率。
Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages

- 通过辅助翻译和语音嵌入构建语义锚点,指导解码器生成
- 在两个30小时的闽南语、客家话数据集上超越现有方法
- 只需小型语音转文本模型即可自动生成有效语义锚点
低资源自动语音识别因标注语料稀缺,难以提供足够的监督信号。为此,我们提出SAMA-ASR,一种轻量级适配机制,在解码器中引入来自辅助翻译的语义锚点和来自语音的声学锚点;该机制可应用于类似编码器-解码器多任务语音模型。通过跨模态适配,SAMA-ASR将翻译生成的语义嵌入与语音嵌入结合,在词元预测前融合整体语义与语音依据。评估时,语义锚点可由上游语音转文本翻译器自动生成,无需人工提供。在涵盖两种低资源汉语方言(台湾闽南语、客家话)的两个30小时数据集上的实验表明,SAMA-ASR优于仅使用声学信号、基于提示或仅依赖语义翻译的基线方法,并在实际自动化语义锚设置下依然有效;译码器容量分析显示,紧凑的语音转文本模型即可生成有用语义锚。
原文摘要 · Abstract (English)
Low-resource ASR remains difficult because scarce transcripts provide limited supervised evidence for target-side generation. To address this gap, we propose SAMA-ASR, a lightweight adapter mechanism that augments the decoder with semantic anchors from auxiliary translations and an acoustic anchor from speech; in principle, the mechanism can be applied to similar encoder--decoder multitask speech models. Through cross-modal adaptation, SAMA-ASR conditions decoder states on translation-derived semantic embeddings and a speech embedding, combining utterance-level meaning with speech-grounded evidence before token prediction. At evaluation time, these semantic anchors can be generated automatically by an upstream speech-to-text translator rather than supplied as oracle translations. Experiments on two 30-hour datasets covering the low-resource Sinitic varieties Taiwanese Hokkien and Hakka show that SAMA-ASR improves over acoustic, prior prompt-based, and semantic-only translation-guided baselines and remains effective in practical automatic semantic-anchor settings; translator-capacity analyses show that useful semantic anchors can be produced by a compact ST model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。