arXiv:2607.10790eess.AS2026-07

用语音化技术生成更贴近真实口语的L2英语发音,提升自动评分效果。

Data Augmentation for L2 English Speaking Assessment using TTS

论文配图:Data Augmentation for L2 English Speaking Assessment using TTS
图 1 · 摘自论文原文
  • 将书面语转为口语风格文本再合成语音,缩小写与说的差异
  • 按水平匹配说话人使合成语音最稳定,评分准确率显著提升
  • 适合做语言测评数据增强的研究者和教育技术开发者

自动评估第二语言(L2)口语能力依赖大规模标注语音数据,但远少于广泛可用的书面语学习语料。一个有前景的方向是利用文本到语音(TTS)和语音克隆技术,将书面L2产出转化为合成语音。然而,书面语与口语存在根本差异:口语包含不流畅和话语标记,而书面语更计划性强、结构复杂。因此需明确生成适用于评估的合成L2语音所需条件。我们通过使用公开的COREFL语料库(包含相同学习者在相同问题下跨模态的配对口语与书面回答),系统分析了说话人与文本的关系。提出框架:首先使用大语言模型将书面语转换为口语风格转录文本(“speechification”),再通过TTS/语音克隆模型生成语音。为分配语音,研究了基于学习者属性(水平、母语、两者或均无)的说话人-文本匹配策略。在语言评估任务中验证数据增强方法,在wav2vec2(音频基础)与ModernBERT(文本基础)评分系统中均表现提升。结果表明,按水平匹配说话人与文本的策略最鲁棒;原始书面文本与口语存在显著偏差,而speechification显著缩小差距并提升评分性能。

原文摘要 · Abstract (English)

Automated assessment of second language (L2) speaking proficiency relies on large-scale annotated speech data, which remains scarce compared to widely available written learner corpora. A promising direction for addressing this imbalance is to use text-to-speech (TTS) and voice cloning to convert written L2 production into synthetic speech. However, written and spoken L2 differ fundamentally: spontaneous speech includes disfluencies and discourse markers, while writing is more planned and complex. This raises the question of what is required to generate synthetic L2 speech suitable for assessment. We address this through a systematic analysis of speaker-text relationships using COREFL, a publicly available corpus containing paired spoken and written responses from the same L2 learners to the same questions across modalities. In our proposed framework, we first address the structural differences between written and spoken language by transforming written responses into spoken-style transcripts ("speechification") using a large language model. These transcripts are then converted into speech using a TTS/voice-cloning model. To assign a voice to each synthetic response, we investigate different speaker-text pairing strategies based on shared learner attributes (proficiency level, first language, both, or neither). We evaluate our data augmentation techniques on the language assessment task, with improvements shown in both wav2vec2 (audio-based) and ModernBERT (text-based) scoring systems. Results show that matching speakers and texts by proficiency level yields the most robust synthetic speech. Moreover, raw written text leads to a strong mismatch with spoken language, while speechification substantially reduces this gap and improves grading performance.

语音生成数据增强L2评估TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。