区分孟加拉语语音识别中的重复错误与语法重叠,构建首个公开语料库。
Distinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis
- 构建2万条标注语料,区分语音识别中的无意重复与有意语法重叠。
- 基于孟加拉语BERT微调模型达84.78%准确率,显著优于大模型少样本提示。
- 为低资源语言文本规范化提供可复用的基准数据与方法框架。
自动语音识别(ASR)转录本在低资源语言如孟加拉语中存在关键歧义:词-词重复可能是无意的停顿错误(重复不流畅性),也可能是有意的语法构词现象(形态重叠)。标准去噪方法会误删有效语言信息。为此,我们首次构建了公开可用的20,000行孟加拉语语料库,人工标注明确区分两种现象。我们在两种范式下进行基准测试:前沿多语言大语言模型(LLMs)与任务专用编码器模型微调。LLMs通过少样本提示达到最高82.68%准确率;而微调后的语言特异性孟加拉语BERT模型表现最优,准确率达84.78%,F1得分为0.677。该工作建立了一个强健的、具有语言学依据的基准,为开发保留语义的孟加拉语文本规范化系统提供了必要数据支持。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) transcripts, especially in low-resource languages like Bangla, contain a critical ambiguity: word-word repetitions can be either Repetition Disfluency (unintentional ASR error/hesitation) or Morphological Reduplication (a deliberate grammatical construct). Standard disfluency correction fails by erroneously deleting valid linguistic information. To solve this, we introduce the first publicly available, 20,000-row Bangla corpus, manually annotated to explicitly distinguish between these two phenomena in noisy ASR transcripts. We benchmark this novel resource using two paradigms: state-of-the-art multilingual Large Language Models (LLMs) and task-specific fine-tuning of encoder models. LLMs achieve competitive performance (up to 82.68\% accuracy) with few-shot prompting. However, fine-tuning proves superior, with the language-specific BanglaBERT model achieving the highest accuracy of 84.78\% and an F1 score of 0.677. This establishes a strong, linguistically-informed baseline and provides essential data for developing sophisticated, semantic-preserving text normalization systems for Bangla.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。