arXiv:2608.13101cs.CLeess.AS2026-08

用语音模型与大模型结合,让口语评分更准且可解释。

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

论文配图:CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model
图 1 · 摘自论文原文
  • 分两步:先提取语音内容和发音特征,再由大模型综合判断
  • 在新数据集上误差仅0.358,比之前最好结果还低一半参数量
  • 适合需要透明评分的教育场景,尤其关注表达与内容分离

自动口语评估研究越来越多地采用多模态语音大模型来评估学习者的口语表现。然而,现有研究对声学信息与内容信息在预测中的贡献程度及其稳定性分析有限。本文提出CASA,一种结合Whisper-medium与Qwen3.5-2B的简化架构,在保持高性能的同时实现了语音表达与内容理解的更可解释分离。在Speak & Improve Corpus 2025上,CASA取得0.358的均方根误差(RMSE),优于先前最佳结果,且推理参数量约为其一半。该通用架构无需结构修改即可适配其他口语评估数据集,依赖三个手工设计的流利度特征。通过消融实验与重复运行,我们分析了声学与内容信息的独立及互补贡献,考察了性能波动性,并验证了大语言模型推理在无训练内容校验中的潜力。

原文摘要 · Abstract (English)

Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.

口语评估语音大模型可解释性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。