arXiv:2510.16387cs.CLcs.AI2025-10中稿 · ICASSP 2026被引 1

不微调也能用,Whisper自动识别二语口语水平

Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment

  • 从Whisper中间层提取声学与语言特征,轻量级分类器判断口语能力
  • 在GEPT数据集上超越现有顶尖方法,含多模态模型
  • 无需任务微调即可捕捉发音流利度与语义信息,适合教育评估场景

本文探索了成熟语音识别模型Whisper在第二语言口语评估(SLA)中的潜在能力。不同于以往仅分析其输出转录文本的研究,本方法通过提取Whisper隐藏表示中的声学与语言特征,深入挖掘其内在表征。仅在Whisper的中间和最终输出上训练一个轻量级分类器,该方法在GEPT图片描述数据集上表现优异,超越现有先进基线,包括多模态方法。进一步引入图像与文本提示作为辅助相关性线索,实现性能提升。深入分析揭示:即使未进行特定任务微调,Whisper的嵌入已天然编码了口语流利度等级与语义内容,凸显其作为口语理解基础模型的巨大潜力。

原文摘要 · Abstract (English)

In this paper, we explore the untapped potential of Whisper, a well-established automatic speech recognition (ASR) foundation model, in the context of L2 spoken language assessment (SLA). Unlike prior studies that extrinsically analyze transcriptions produced by Whisper, our approach goes a step further to probe its latent capabilities by extracting acoustic and linguistic features from hidden representations. With only a lightweight classifier being trained on top of Whisper's intermediate and final outputs, our method achieves strong performance on the GEPT picture-description dataset, outperforming existing cutting-edge baselines, including a multimodal approach. Furthermore, by incorporating image and text-prompt information as auxiliary relevance cues, we demonstrate additional performance gains. Finally, we conduct an in-depth analysis of Whisper's embeddings, which reveals that, even without task-specific fine-tuning, the model intrinsically encodes both ordinal proficiency patterns and semantic aspects of speech, highlighting its potential as a powerful foundation for SLA and other spoken language understanding tasks.

语音评估Whisper二语学习基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。