arXiv:2508.12591cs.CLcs.AI2025-08中稿 · IEEE ASRU 2025

用多模态大模型提升口语评估,首次系统验证其全面优势

Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum Learning

  • 提出语音优先的多模态训练策略,先建语音基础再融合文本
  • 在基准数据集上将整体评估相关性从0.783提升至0.846
  • 对发音表现评估精度提升4%,适合教育AI与自动评分研发者

传统自动口语评估系统存在模态局限:纯文本方法缺失声学信息,纯音频方法忽略语义上下文。多模态大语言模型(MLLM)为全面评估提供了前所未有的机遇,可在统一框架中同时处理音频与文本。本文首次系统研究了MLLM在全面口语评估中的应用,证实其在内容与语言使用方面的优越性能。然而,在发音表现评估方面仍面临独特挑战,需专用训练策略。为此,我们提出语音优先多模态训练(SFMT),基于课程学习原则,在跨模态融合前建立更稳健的语音建模基础。在基准数据集上的一系列实验表明,基于MLLM的系统可将整体评估相关性(PCC)从0.783提升至0.846。尤其在发音表现评估上,SFMT相较传统方法绝对提升4%准确率,为自动口语评估开辟新路径。

原文摘要 · Abstract (English)

Traditional Automated Speaking Assessment (ASA) systems exhibit inherent modality limitations: text-based approaches lack acoustic information while audio-based methods miss semantic context. Multimodal Large Language Models (MLLM) offer unprecedented opportunities for comprehensive ASA by simultaneously processing audio and text within unified frameworks. This paper presents a very first systematic study of MLLM for comprehensive ASA, demonstrating the superior performance of MLLM across the aspects of content and language use . However, assessment on the delivery aspect reveals unique challenges, which is deemed to require specialized training strategies. We thus propose Speech-First Multimodal Training (SFMT), leveraging a curriculum learning principle to establish more robust modeling foundations of speech before cross-modal synergetic fusion. A series of experiments on a benchmark dataset show MLLM-based systems can elevate the holistic assessment performance from a PCC value of 0.783 to 0.846. In particular, SFMT excels in the evaluation of the delivery aspect, achieving an absolute accuracy improvement of 4% over conventional training approaches, which also paves a new avenue for ASA.

口语评估多模态大模型课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。