用语音嵌入追踪听障儿童语音向成人靠拢的过程。
Self-Supervised Speech Representations Track Spoken Language Convergence to Adult Models in Infants and Children Who Are Deaf/Hard-of-Hearing

- 用HuBERT提取儿童日常语音的嵌入表示,直接分析语音演变。
- 听觉年龄越大,儿童与成人的语音嵌入距离越小,体现语音趋同。
- 该方法可替代人工标注,适合跨语言、大规模发展评估。
语言发展表现为儿童语音逐渐趋近成人模式。传统测量依赖详细转录和语言专业知识,难以跨语言和人群扩展。本文利用语音嵌入,从儿童日常生活的长时录音中直接捕捉这一趋同过程。基于HuBERT-BASE模型,对听障儿童及其女性成年照料者(累计观察超925小时)的语音发声提取嵌入表示。控制音高和发声长度后,发现儿童与照料者之间的嵌入距离随听觉年龄增长而减小,表明儿童语音模式随发育逐渐向成人靠拢。该单一距离指标还与婴幼儿至学龄前阶段的多项标准化言语和语言测验结果相关。研究结果为从儿童日常生活语音中实现可扩展、无语言依赖的语言发展评估提供了新路径。
原文摘要 · Abstract (English)
Language development is characterized by a gradual convergence of children's speech toward adult patterns. Measuring this process has traditionally required detailed transcription and language-specific expertise, limiting scalability across languages and populations. Here, we use speech embeddings to capture this convergence directly from the acoustic signal in longform, child-centered recordings, taken as children go about their daily lives. Using HuBERT-BASE, we extracted embeddings from speech vocalizations of children who are deaf/hard-of-hearing and their female adult caregivers ($>$925 hrs. observation). Embedding distance between children and caregivers decreased with hearing age, controlling for pitch and vocalization length, indicating, as expected, that children's speech patterns converge to caregivers over development. This single distance metric likewise related to multiple standardized measures of speech and language from infancy through preschoolhood. These results suggest a path toward scalable, language-neutral assessment of spoken language development from children's everyday lives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。