arXiv:2506.05121cs.CLcs.SD2025-06被引 3

融合声学与语义模型,提升口语能力自动评估精度。

The NTNU System at the S&I Challenge 2025 SLA Open Track

  • 用W2V提取语音特征,Phi-4 MLLM理解语义,融合评分
  • 在比赛中取得RMSE 0.375,仅次于第一名的0.364
  • 适合需要兼顾语音与语义评估的研究者和开发者

近年来,语音语言评估(SLA)研究采用BERT和wav2vec 2.0(W2V)等神经模型,从语言和声学双模态评估口语能力。尽管两者均能捕捉口语表现相关特征,但各有局限:基于BERT的方法依赖ASR转录文本,难以反映语调与发音线索;而基于W2V的方法虽擅长建模声学特征,却缺乏语义可解释性。为此,我们提出一种将W2V与Phi-4多模态大语言模型(MLLM)通过评分融合策略相结合的系统。该系统在Speak & Improve Challenge 2025官方测试集上取得0.375的均方根误差(RMSE),排名第二。相较之下,第一名、第三名及官方基线系统的RMSE分别为0.364、0.384和0.444。

原文摘要 · Abstract (English)

A recent line of research on spoken language assessment (SLA) employs neural models such as BERT and wav2vec 2.0 (W2V) to evaluate speaking proficiency across linguistic and acoustic modalities. Although both models effectively capture features relevant to oral competence, each exhibits modality-specific limitations. BERT-based methods rely on ASR transcripts, which often fail to capture prosodic and phonetic cues for SLA. In contrast, W2V-based methods excel at modeling acoustic features but lack semantic interpretability. To overcome these limitations, we propose a system that integrates W2V with Phi-4 multimodal large language model (MLLM) through a score fusion strategy. The proposed system achieves a root mean square error (RMSE) of 0.375 on the official test set of the Speak & Improve Challenge 2025, securing second place in the competition. For comparison, the RMSEs of the top-ranked, third-ranked, and official baseline systems are 0.364, 0.384, and 0.444, respectively.

语音评估多模态评分融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。