构建27个儿童语音数据集并嵌入伦理治理,解决儿童语言研究中的数据与隐私难题。
Deriving Benchmarking Datasets from Long-Form Recordings: Challenges and Opportunities

- 统一收集27个儿童语音数据集,使用开源工具确保格式一致
- 建立可复现的语音处理基准测试流程,支持跨语言评估
- 引入角色化伦理系统ELSI,实现机器学习流程中的隐私保护
儿童中心的长时录音(LFRs)是研究早期语言发展的生态有效数据源,但存在三大限制:第一,不同站点采集的数据格式和同意机制不统一,导致跨库使用困难;第二,缺乏标准化基准,难以评估工具在不同语言和条件下的泛化能力;第三,机器学习工作流常忽视敏感儿童语音的隐私约束。本文提出一个框架,涵盖三方面解决方案:(S1)使用开源工具构建27个儿童中心数据集的标准化集合;(S2)开发可复现的四种语音处理基准测试流程;(S3)提出ELSI——一种嵌入伦理治理的角色化生态系统。通过语音类型分类案例研究验证了该框架,并表明三个方案相互依赖、缺一不可。
原文摘要 · Abstract (English)
Long-form recordings (LFRs) of child-centered audio are ecologically valid sources for studying early language development, but three problems limit their use. First, LFR corpora are collected across sites with heterogeneous formats and consent structures, making cross-corpus use non-trivial. Second, without standardized benchmarks, assessing whether tools generalize across languages and conditions is hard. Third, ML workflows rarely respect privacy constraints governing sensitive child speech. This paper presents a framework addressing all three: a standardized collection of 27 child-centered datasets built with open-source tools (S1); a replicable pipeline for four speech-processing benchmarks (S2); and ELSI, a role-based ecosystem embedding ethical governance into the ML workflow (S3). We demonstrate the framework via a voice type classification case study and show the three solutions are mutually dependent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。