用多模型共识修正儿童语音数据的时间戳,提升标注精度。
CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

- 通过多个语音识别模型对齐并投票生成更准的时间边界
- 构建413小时儿童语音数据集,283小时高质量子集可直接用于训练
- 在4个外部数据集上平均降低19.5%的词错误率,适合语音模型训练
CHILDES是一个大规模儿童语音语料库,包含自然对话中的长时录音,是研究儿童语言发展的宝贵资源。但其中提供的语音片段时间戳常存在噪声、缺失或错位,导致难以准确定位语音内容,限制了其在语音模型训练与评估中的直接应用。本文提出BEACON(边界估计通过对齐共识),一种基于多模型时间戳集成的校正框架。该方法首先将各现成语音识别模型的词级时间戳与人工转录文本对齐,再通过共识投票确定最终的语音片段边界。该框架不依赖特定语料库,适用于任意配有可信转录但时间戳不可靠或缺失的长音频数据,提供通用的时间戳修复方案。基于此流程,我们构建并发布了一个413小时的通用儿童语音数据集,其中包含283小时经质量控制的子集,可用于语音识别训练。在该子集上微调模型,可在四个域外儿童语音基准上实现平均19.5%的相对词错误率降低。
原文摘要 · Abstract (English)
CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and language development. However, utterance-level timestamps provided in this corpus are often noisy, incomplete, or misaligned with the audio. As a result, utterances cannot always be reliably localized within long recordings, which limits the direct use of these data for training and evaluating speech models. In this work, we propose BEACON (Boundary Estimation via Alignment CONsensus), an ensemble timestamp-curation framework that refines utterance-level timestamps by aggregating knowledge from multiple off-the-shelf ASR models. Specifically, each model's word-level timestamp predictions are first aligned to provided human transcripts, and the final utterance time boundaries are determined by a consensus voting strategy. The framework is corpus-agnostic and applies to any long-form recording paired with a trusted transcript whose timestamps are unreliable or missing, offering a general recipe for timestamp curation. Leveraging this pipeline, we curate and release a 413-hour general-purpose child-speech dataset with corrected utterance-level timestamps, together with a 283-hour quality-controlled subset for ASR training. Fine-tuning on this subset yields up to an average 19.5% relative WER reduction on four out-of-domain child-speech benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。