arXiv:2503.11229cs.SDcs.CL2025-03被引 8

用GPT-4o评估发音,能生成实用反馈并精准打分。

Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment

  • 用GPT-4o处理语音,多粒度分析发音问题。
  • 评分与人工标注相关性高,反馈质量由LLM评估达标。
  • 适合教育科技、智能评测系统开发者参考。

大型多模态模型(LMMs)在多个领域表现优异。本文探索其在发音评估任务中的潜力,重点评估生成式预训练变换器(GPT)模型——GPT-4o的能力。研究考察其在不同粒度和维度下处理语音与音频进行发音评估的性能,尤其关注反馈生成与评分效果。实验基于公开数据集Speechocean762,评估包含两方面:多层级评分准确性与生成反馈的实用性。评分结果与原始数据集中的人工标注分数对比,反馈质量则通过大型语言模型(LLMs)评估。结果表明,将LMMs与传统方法结合可有效提升发音评估能力,揭示了模型优势并指明改进方向。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have demonstrated exceptional performance across a wide range of domains. This paper explores their potential in pronunciation assessment tasks, with a particular focus on evaluating the capabilities of the Generative Pre-trained Transformer (GPT) model, specifically GPT-4o. Our study investigates its ability to process speech and audio for pronunciation assessment across multiple levels of granularity and dimensions, with an emphasis on feedback generation and scoring. For our experiments, we use the publicly available Speechocean762 dataset. The evaluation focuses on two key aspects: multi-level scoring and the practicality of the generated feedback. Scoring results are compared against the manual scores provided in the Speechocean762 dataset, while feedback quality is assessed using Large Language Models (LLMs). The findings highlight the effectiveness of integrating LMMs with traditional methods for pronunciation assessment, offering insights into the model's strengths and identifying areas for further improvement.

发音评估GPT-4o多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。