arXiv:2505.16168cs.SDeess.AS2025-05中稿 · INTERSPEECH 2025被引 1

根据语音难易度选择性调用ASR模型,降错率18.7%、成本减半

Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty

  • 基于语音大模型判断输入简单与否,决定是否调用高精度ASR
  • 相比语音大模型错率降低18.7%,调用成本仅为传统方法一半
  • 适合资源受限但需多语言支持的实时语音识别场景

尽管多语言自动语音识别(ASR)系统已取得显著进展,能用单一模型处理多种语言,但固有的语言差异和数据不平衡仍限制其在所有语言上的表现。现有方法依赖语言识别(LID)模型将语音路由至对应ASR模型,但调用高性能商用模型成本高昂,且易因误分类导致错误。为此,我们提出SIMA,一种面向多语言ASR的选择性调用方法,可根据输入语音的难度动态决策。该方法基于语音大语言模型(SLLM),评估输入是否足够简单可直接转录,或需调用先进ASR模型。实验表明,与SLLM相比,SIMA将词错误率降低18.7%,调用成本较基于LID的方法减少50%。在三个数据集上的测试验证了其可扩展性和成本效益,适用于多语言语音识别应用。

原文摘要 · Abstract (English)

Although multilingual automatic speech recognition (ASR) systems have significantly advanced, enabling a single model to handle multiple languages, inherent linguistic differences and data imbalances challenge SOTA performance across all languages. While language identification (LID) models can route speech to the appropriate ASR model, they incur high costs from invoking SOTA commercial models and suffer from inaccuracies due to misclassification. To overcome these, we propose SIMA, a selective invocation for multilingual ASR that adapts to the difficulty level of the input speech. Built on a spoken large language model (SLLM), SIMA evaluates whether the input is simple enough for direct transcription or requires the invocation of a SOTA ASR model. Our approach reduces word error rates by 18.7% compared to the SLLM and halves invocation costs compared to LID-based methods. Tests on three datasets show that SIMA is a scalable, cost-effective solution for multilingual ASR applications.

多语言ASR成本优化语音大模型选择性调用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。