arXiv:2506.11747cs.SDcs.LG2025-06被引 1

用智能筛选提升儿童语音转录准确率,让大规模语言研究成为可能。

Enabling automatic transcription of child-centered audio recordings from real-world environments

  • 先筛选出适合自动识别的语音片段,再用ASR处理,避免噪声干扰
  • 可转录13%语音,平均错误率仅18%,远低于全量处理的51%
  • 结果与人工标注高度一致,适合做儿童语言发展研究

儿童佩戴麦克风获取的长时间音频记录已成为研究儿童语言经验及其对语言发展影响的标准方法。若能对长时语音进行自动转录,将支持多层级语言分析,但典型语料规模庞大,难以全面人工标注。同时,由于真实环境音频噪声大、无约束,现有自动语音识别(ASR)系统难以有效应用。以往研究假设必须对整段录音进行处理,本工作提出新方法:自动识别出可用现代ASR可靠转录的语音片段,实现对典型长时数据中可观比例语音的自动且较准确转录。在四个英语长时音频语料上验证,该方法在转录总语音量13%的情况下,达到中位词错误率(WER)0%、均值WER 18%;而全量转录无筛选时,中位WER为52%、均值51%。对比自动转录与人工标注的词频分布,相关系数分别为r=0.92(所有词)和r=0.98(出现≥5次的词)。该研究为儿童中心长时音频的自动化语言分析提供了切实可行的路径。

原文摘要 · Abstract (English)

Longform audio recordings obtained with microphones worn by children-also known as child-centered daylong recordings-have become a standard method for studying children's language experiences and their impact on subsequent language development. Transcripts of longform speech audio would enable rich analyses at various linguistic levels, yet the massive scale of typical longform corpora prohibits comprehensive manual annotation. At the same time, automatic speech recognition (ASR)-based transcription faces significant challenges due to the noisy, unconstrained nature of real-world audio, and no existing study has successfully applied ASR to transcribe such data. However, previous attempts have assumed that ASR must process each longform recording in its entirety. In this work, we present an approach to automatically detect those utterances in longform audio that can be reliably transcribed with modern ASR systems, allowing automatic and relatively accurate transcription of a notable proportion of all speech in typical longform data. We validate the approach on four English longform audio corpora, showing that it achieves a median word error rate (WER) of 0% and a mean WER of 18% when transcribing 13% of the total speech in the dataset. In contrast, transcribing all speech without any filtering yields a median WER of 52% and a mean WER of 51%. We also compare word log-frequencies derived from the automatic transcripts with those from manual annotations and show that the frequencies correlate at r = 0.92 (Pearson) for all transcribed words and r = 0.98 for words that appear at least five times in the automatic transcripts. Overall, the work provides a concrete step toward increasingly detailed automated linguistic analyses of child-centered longform audio.

语音转录儿童语言自动标注自然语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。