测试SLAM-ASR在不同语音场景下的表现,发现其泛化能力差。
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
- 用线性连接器将语音编码器与大模型结合,构建简单ASR系统
- 跨领域和语音扰动下性能显著下降,噪声使识别率大幅降低
- 为语音大模型部署提供调参建议,适合关注鲁棒性的研究者
近期研究显示,通过线性连接语音基础编码器与大语言模型(LLMs),可实现强大的自动语音识别(ASR)能力。尽管结果令人印象深刻,但这类简单方法在不同场景和语音条件下的鲁棒性仍不明确,例如领域迁移和语音扰动。本文通过一系列消融实验,评估了一种近期广泛采用的SLAM-ASR方法。我们发现,该架构在跨领域设置中表现不佳;即使在同领域数据上,语音速率变化或添加噪声等扰动也会显著降低性能。这些实证发现为基于LLM的ASR模型的微调与配置提供了关键指导,有助于根据数据特征和计算资源设计更稳健的系统。
原文摘要 · Abstract (English)
Recent research has demonstrated that training a linear connector between speech foundation encoders and large language models (LLMs) enables this architecture to achieve strong ASR capabilities. Despite the impressive results, it remains unclear whether these simple approaches are robust enough across different scenarios and speech conditions, such as domain shifts and speech perturbations. In this paper, we address these questions by conducting various ablation experiments using a recent and widely adopted approach called SLAM-ASR. We present novel empirical findings that offer insights on how to effectively utilize the SLAM-ASR architecture across a wide range of settings. Our main findings indicate that SLAM-ASR exhibits poor performance in cross-domain evaluation settings. Additionally, speech perturbations on in-domain data, such as changes in speech rate or additive noise, can significantly degrade performance. Our findings offer critical insights for fine-tuning and configuring robust LLM-based ASR models, tailored to different data characteristics and computational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。