arXiv:2605.29257cs.SD2026-05

构建儿童发声全周期音频基准,助力语言发展评估

ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood

论文配图:ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood
图 1 · 摘自论文原文
  • 覆盖从出生到学龄期的多种儿童声音信号,整合17个数据集
  • 在20+子任务上验证模型性能,识别儿童语音信号准确率显著提升
  • 适合儿童语言发育研究、智能教育与临床评估领域使用

我们提出ChildVox,一个用于刻画儿童沟通中多样化声学信号的新基准。该基准覆盖从出生到学龄期的完整发育轨迹,涵盖生理声音、非语言发声、典型音节及口语表达。ChildVox整合了17个以儿童为中心的音频与语音数据集,包含20多个子任务,支持跨数据集、跨领域的系统性比较。我们在生理声音分类、发声与音节建模、语音质量评估与识别等任务上,对自监督、语音识别导向及大型音频-语言模型进行了评估。结果显示,ChildVox能有效识别儿童广泛声学信号,支持语言水平表征与年龄相关语音产出追踪等下游应用。

原文摘要 · Abstract (English)

We present ChildVox, a novel benchmark for characterizing the diverse acoustic signals through which children communicate. Specifically, ChildVox follows the full developmental trajectory from birth through school age, covering physiological sounds, non-linguistic vocalizations, canonical syllables, and spoken language. ChildVox integrates more than 20 sub-tasks across 17 child-centered audio and speech datasets, enabling systematic cross-corpus and cross-domain comparison. We evaluate a representative range of audio and speech foundation models, including self-supervised, ASR-oriented, and large audio-language models, on tasks including physiological sound classification, vocalization and canonical syllables modeling, and speech quality assessment and recognition. Benchmark results show that ChildVox provides a suite of high-performance models in recognizing a wide range of acoustic signals from children, supporting downstream applications such as characterizing children's language levels and tracking speech production with age.

语音分析儿童语言基准测试音频模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。