arXiv:2507.19374cs.CLcs.AI2025-07中稿 · ISCA SLaTE 2025被引 2

自动生成带语法错误的语音文本对,提升口语纠错数据质量。

Data Augmentation for Spoken Grammatical Error Correction

  • 全自动生成含语法错误和口误的语音文本对。
  • 新数据集在保持语言水平评分不变的前提下丰富了错误类型。
  • 适用于提升英语学习者口语纠错模型训练效果。

尽管语法纠错(GEC)已有强大基准数据集,但高质量标注的口语语法纠错(SGEC)数据仍严重不足。本文提出一种完全自动化的音频-文本对生成方法,可合成带有语法错误和口误的语音文本数据。同时设计一系列客观评估指标,用于衡量生成数据质量并选择最优数据集。目标是在保留原始数据文本与声学特征的基础上,引入新型错误类型,从而扩充并丰富原始语料库,且不改变第二语言学习者的语言评估分数。实验在首个公开的带语法错误标注的语音数据集S&I Corpus上进行,验证了该增强语料在书面GEC(仅文本部分)与口语GEC(音视频-文本对)任务中的有效性。

原文摘要 · Abstract (English)

While there exist strong benchmark datasets for grammatical error correction (GEC), high-quality annotated spoken datasets for Spoken GEC (SGEC) are still under-resourced. In this paper, we propose a fully automated method to generate audio-text pairs with grammatical errors and disfluencies. Moreover, we propose a series of objective metrics that can be used to evaluate the generated data and choose the more suitable dataset for SGEC. The goal is to generate an augmented dataset that maintains the textual and acoustic characteristics of the original data while providing new types of errors. This augmented dataset should augment and enrich the original corpus without altering the language assessment scores of the second language (L2) learners. We evaluate the use of the augmented corpus both for written GEC (the text part) and for SGEC (the audio-text pairs). Our experiments are conducted on the S\&I Corpus, the first publicly available speech dataset with grammar error annotations.

口语纠错数据增强语音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。