针对语音识别模型在特定领域表现下降的问题,提出一套以指标为导向的微调框架。
Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
- 基于性能指标优化学习率,结合领域数据变换与增强。
- 在Whisper、Qwen2-Audio等模型上验证,显著提升跨领域识别准确率。
- 适合需要快速适配新场景的语音系统研发人员使用。
自动语音识别(ASR)模型在通用场景下已取得卓越精度,但在特定领域应用中常因数据不匹配和语言差异导致性能下降。这一问题在基于大语言模型(LLM)的ASR系统中尤为突出,其庞大规模与复杂的训练动态使得有效微调变得困难。为此,本文提出一种原则性且以指标为导向的微调框架,用于适应传统及LLM-based ASR模型至特定领域。该框架强调基于性能指标的学习率优化,并结合领域特定的数据变换与增强策略。我们在多个领域、多语言、多长度的数据集上对Whisper、Whisper-Turbo和Qwen2-Audio等先进模型进行了实证评估。结果不仅验证了所提框架的有效性,还建立了提升领域特定ASR性能并防止过拟合的实用协议。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) models have achieved remarkable accuracy in general settings, yet their performance often degrades in domain-specific applications due to data mismatch and linguistic variability. This challenge is amplified for modern Large Language Model (LLM)-based ASR systems, whose massive scale and complex training dynamics make effective fine-tuning non-trivial. To address this gap, this paper proposes a principled and metric-driven fine-tuning framework for adapting both traditional and LLM-based ASR models to specialized domains. The framework emphasizes learning rate optimization based on performance metrics, combined with domain-specific data transformation and augmentation. We empirically evaluate our framework on state-of-the-art models, including Whisper, Whisper-Turbo, and Qwen2-Audio, across multi-domain, multilingual, and multi-length datasets. Our results not only validate the proposed framework but also establish practical protocols for improving domain-specific ASR performance while preventing overfitting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。