arXiv:2505.21137cs.CLcs.SD2025-05被引 3

用伪标签扩增数据,提升语音语法纠错与反馈效果

Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction

  • 通过伪标签将训练数据从77小时扩至2500小时
  • 提示词使纠错准确率微升,反馈生成显著改善
  • 大模型需提示词才能发挥伪标签优势,适合语言学习场景

语音语法纠错(SGEC)与反馈(SGECF)对第二语言学习者、教师和考生至关重要。传统方法采用包含语音识别(ASR)、不流畅检测与去除、语法纠错(GEC)的级联流程。随着端到端(E2E)语音基础模型的发展,本文研究其在SGEC与反馈生成中的有效性。提出一种伪标签流程,将训练数据从77小时扩展至约2500小时,显著提升性能。此外,使用流利转录作为提示,对基于Whisper的E2E SGEC模型进行提示,纠错表现略有提升,反馈生成效果更明显。最后评估模型规模影响:在大Whisper模型上,伪标签不再带来增益,但提示训练仍具优势。

原文摘要 · Abstract (English)

Spoken Grammatical Error Correction (SGEC) and Feedback (SGECF) are crucial for second language learners, teachers and test takers. Traditional SGEC systems rely on a cascaded pipeline consisting of an ASR, a module for disfluency detection (DD) and removal and one for GEC. With the rise of end-to-end (E2E) speech foundation models, we investigate their effectiveness in SGEC and feedback generation. This work introduces a pseudo-labelling process to address the challenge of limited labelled data, expanding the training data size from 77 hours to approximately 2500 hours, leading to improved performance. Additionally, we prompt an E2E Whisper-based SGEC model with fluent transcriptions, showing a slight improvement in SGEC performance, with more significant gains in feedback generation. Finally, we assess the impact of increasing model size, revealing that while pseudo-labelled data does not yield performance gain for a larger Whisper model, training with prompts proves beneficial.

语音纠错伪标签提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。