无需真实标签,用多模态嵌入预测语音识别准确率。
On the Robust Approximation of ASR Metrics
- 用语音与文本的统一嵌入空间+代理模型生成指标代理值。
- 在14个数据集上对40多个模型预测,误差低于一位数。
- 适合无标注数据时快速评估语音模型性能,节省人力成本。
近年来,语音基础模型的发展主要依赖于模型规模和数据量的扩大,使其能够完成包括语音识别在内的多种任务。传统上,语音识别模型通过词错误率(WER)和字符错误率(CER)等指标进行评估,这些指标依赖于真实标签。然而,由于来自多样化领域和测试条件的标注数据有限,模型在标准基准之外的真实泛化能力尚不明确;且数据标注成本高、耗时长。为解决此问题,我们提出一种新颖的无标签方法来近似语音识别性能指标,无需真实标签。该方法利用语音与转录文本在统一空间中的多模态嵌入,并结合高质量代理模型计算代理指标。这些特征被用于训练回归模型,以预测关键的语音识别指标,如词错误率(WER)和字符错误率(CER)。我们在14个数据集上对超过40个模型进行了实验,涵盖标准和真实场景下的测试条件。结果表明,在所有实验配置中,我们的方法对指标的近似误差均在一位数以内,相比最新基线提升超过50%。
原文摘要 · Abstract (English)
Recent advances in speech foundation models are largely driven by scaling both model size and data, enabling them to perform a wide range of tasks, including speech recognition. Traditionally, ASR models are evaluated using metrics like Word Error Rate (WER) and Character Error Rate (CER), which depend on ground truth labels. As a result of limited labeled data from diverse domains and testing conditions, the true generalization capabilities of these models beyond standard benchmarks remain unclear. Moreover, labeling data is both costly and time-consuming. To address this, we propose a novel label-free approach for approximating ASR performance metrics, eliminating the need for ground truth labels. Our method utilizes multimodal embeddings in a unified space for speech and transcription representations, combined with a high-quality proxy model to compute proxy metrics. These features are used to train a regression model to predict key ASR metrics like Word Error Rate (WER) and Character Error Rate (CER). We experiment with over 40 models across 14 datasets representing both standard and in-the-wild testing conditions. Our results show that we approximate the metrics within a single-digit absolute difference across all experimental configurations, outperforming the most recent baseline by more than 50\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。