arXiv:2510.04463cs.SDeess.AS2025-10中稿 · ASRU 2025

用大模型评估自监督语音模型,无需额外训练。

Evaluating Self-Supervised Speech Models via Text-Based LLMS

  • 输入语音模型的离散令牌和领域提示,让大模型计算对数似然。
  • 评估分数与语音识别任务表现高度相关(相关系数0.92)。
  • 还能生成可用于说话人验证的推理时嵌入。

自监督学习(SSL)因其低成本标注即可获得丰富表征而受到广泛关注,适用于多种下游任务。然而,由于需要额外训练和评估,其下游性能评估仍具挑战性。现有无任务评估方法也需额外训练或超参数调优。本文提出一种基于大语言模型(LLM)的新评估指标:将来自SSL模型的离散令牌序列及少量领域提示输入LLM,计算平均对数似然。该提示引导上下文学习,使评分更可靠,且无需额外训练或调参。实验表明,基于LLM的评分与自动语音识别任务表现高度相关(相关系数0.92)。此外,研究发现LLM不仅可作为SSL评估工具,还能提供对说话人验证任务有用的推理时嵌入。

原文摘要 · Abstract (English)

Self-Supervised Learning (SSL) has gained traction for its ability to learn rich representations with low labeling costs, applicable across diverse downstream tasks. However, assessing the downstream-task performance remains challenging due to the cost of extra training and evaluation. Existing methods for task-agnostic evaluation also require extra training or hyperparameter tuning. We propose a novel evaluation metric using large language models (LLMs). By inputting discrete token sequences and minimal domain cues derived from SSL models into LLMs, we obtain the mean log-likelihood; these cues guide in-context learning, rendering the score more reliable without extra training or hyperparameter tuning. Experimental results show a correlation between LLM-based scores and automatic speech recognition task. Additionally, our findings reveal that LLMs not only functions as an SSL evaluation tools but also provides inference-time embeddings that are useful for speaker verification task.

自监督学习语音模型大模型评估说话人验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。