arXiv:2509.15701cs.CLcs.SD2025-09被引 1

用大模型微调提升发音评估精度,词句级表现佳但音素级仍难。

Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

  • 微调大模型用于多粒度发音评估,提升零样本效果。
  • 词句级评估性能媲美主流系统,音素级仍存挑战。
  • 皮尔逊相关性0.9,斯皮尔曼相关性0.6,反映排序一致性更关键。

自动发音评估(APA)在计算机辅助语言学习中至关重要,需覆盖多个粒度和维度。大型多模态模型(LMMs)为APA带来新机遇,但其在细粒度评估中的有效性尚不明确。本文基于Speechocean762数据集和一个私有语料库,研究LMMs的微调方法。微调显著优于零样本设置,在单粒度任务上达到与公开及商业系统相当的性能。模型在词级和句级表现良好,但音素级评估仍具挑战。我们观察到皮尔逊相关系数(PCC)达0.9,而斯皮尔曼等级相关系数(SCC)约为0.6,表明SCC更能反映排序一致性。这些发现揭示了LMMs在APA中的潜力与局限,并指向未来在细粒度建模与秩感知评估方面的研究方向。

原文摘要 · Abstract (English)

Automatic Pronunciation Assessment (APA) is critical for Computer-Assisted Language Learning (CALL), requiring evaluation across multiple granularities and aspects. Large Multimodal Models (LMMs) present new opportunities for APA, but their effectiveness in fine-grained assessment remains uncertain. This work investigates fine-tuning LMMs for APA using the Speechocean762 dataset and a private corpus. Fine-tuning significantly outperforms zero-shot settings and achieves competitive results on single-granularity tasks compared to public and commercial systems. The model performs well at word and sentence levels, while phoneme-level assessment remains challenging. We also observe that the Pearson Correlation Coefficient (PCC) reaches 0.9, whereas Spearman's rank Correlation Coefficient (SCC) remains around 0.6, suggesting that SCC better reflects ordinal consistency. These findings highlight both the promise and limitations of LMMs for APA and point to future work on fine-grained modeling and rank-aware evaluation.

发音评估大模型微调多模态语音分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。