儿童可穿戴设备语音识别面临数据质量与数量瓶颈
Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
- 三年实验聚焦语音类型分类任务优化
- 特征、模型架构改进仅带来微小性能提升
- 数据相关性与规模更重要,需合规采集共享
通过儿童可穿戴设备收集的语音记录有望彻底改变基础与应用语音科学,实现对儿童自然语言环境与表达的无感捕捉。这一前景依赖于将海量数据转化为可用信息的语音技术。本文总结了三年实验成果,揭示了在提升一项基础任务——语音类型分类性能过程中面临的多重障碍。实验表明,尽管在表示特征、模型架构和参数搜索方面持续优化,性能提升仍十分有限。相比之下,关注数据的相关性与数量带来了更显著进展,凸显了在获得适当授权后开展数据共享的重要性。
原文摘要 · Abstract (English)
Recordings gathered with child-worn devices promised to revolutionize both fundamental and applied speech sciences by allowing the effortless capture of children's naturalistic speech environment and language production. This promise hinges on speech technologies that can transform the sheer mounds of data thus collected into usable information. This paper demonstrates several obstacles blocking progress by summarizing three years' worth of experiments aimed at improving one fundamental task: Voice Type Classification. Our experiments suggest that improvements in representation features, architecture, and parameter search contribute to only marginal gains in performance. More progress is made by focusing on data relevance and quantity, which highlights the importance of collecting data with appropriate permissions to allow sharing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。