arXiv:2506.19315cs.CLcs.AI2025-06中稿 · the ISCA SLaTE-202…被引 1

用Mamba模型联合提升语音评估与错音诊断效果

JCAPT: A Joint Modeling Approach for CAPT

  • 引入Mamba模型结合音韵特征和思维标记,统一建模语音评估与错音诊断
  • 在speechocean762数据集上,错音检测任务性能显著优于已有方法
  • 首次融合音韵归因、状态空间模型与提示技术,适合语音教学系统研发

有效的发音反馈对第二语言学习至关重要,计算机辅助发音训练(CAPT)系统通常包含自动发音评估(APA)和错误发音检测与诊断(MDD)两项关键任务。近期研究表明,联合建模这两项任务可产生互惠效益。本文提出统一框架,采用选择性状态空间模型Mamba,结合音韵特征与思维标记策略,共同增强APA与MDD的可解释性及细粒度时序推理能力。据我们所知,这是首个将音韵归因、基于SSM的建模与提示技术整合于CAPT的研究。在speechocean762基准上的系列实验表明,该模型持续优于先前方法,尤其在MDD任务上表现突出。

原文摘要 · Abstract (English)

Effective pronunciation feedback is critical in second language (L2) learning, for which computer-assisted pronunciation training (CAPT) systems often encompass two key tasks: automatic pronunciation assessment (APA) and mispronunciation detection and diagnosis (MDD). Recent work has shown that joint modeling of these two tasks can yield mutual benefits. Our unified framework leverages Mamba, a selective state space model (SSM), while integrating phonological features and think token strategies to jointly enhance interpretability and fine-grained temporal reasoning in APA and MDD. To our knowledge, this is the first study to combine phonological attribution, SSM-based modeling, and prompting in CAPT. A series of experiments conducted on the speechocean762 benchmark demonstrate that our model consistently outperforms prior methods, particularly on the MDD task.

语音评估Mamba错音检测语言学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。