自动化语音分析中的分类错误会扭曲儿童语言研究结论,该文提出用贝叶斯方法校正偏差。
Classification errors distort findings in automated speech processing: examples and solutions from child-development research
- 采用贝叶斯联合建模,同时分析语音行为与算法误判机制。
- 实测显示,常见语音分类器误差使效应量估计偏移超过20%。
- 提出校准方法可部分恢复真实效应,适合关注数据可信度的研究者。
随着可穿戴录音设备的普及,科研人员越来越多地使用自动化音频分析方法来测量儿童的经历、行为和发育结果,尤其在语言习得研究中广泛使用长时录音。尽管已有大量文献报告主流自动化分类器的准确率与可靠性,但关于分类错误对后续测量与统计推断(如回归分析中的相关性与效应量估计)的下游影响却较少被讨论。本文重点揭示了混淆错误对科学结论的扭曲作用,并提出一种衡量与修正此类误差的方法。具体而言,我们采用贝叶斯方法研究算法误差对关键科学问题的影响,包括兄弟姐妹对儿童语言经验的作用,以及儿童产出与其输入之间的关联。通过在真实与模拟数据上拟合语音行为与算法行为的联合模型,我们发现分类错误会显著扭曲最常用的Lena工具及一个更准确的开源替代方案(ACLEW系统的语音类型分类器)的估计结果。此外,我们证明贝叶斯校准方法可在一定程度上恢复无偏效应量估计,但无法完全杜绝偏差。
原文摘要 · Abstract (English)
With the advent of wearable recorders, scientists are increasingly turning to automated methods of analysis of audio and video data in order to measure children's experience, behavior, and outcomes, with a sizable literature employing long-form audio-recordings to study language acquisition. While numerous articles report on the accuracy and reliability of the most popular automated classifiers, less has been written on the downstream effects of classification errors on measurements and statistical inferences (e.g., the estimate of correlations and effect sizes in regressions). This paper's main contributions are drawing attention to downstream effects of confusion errors, and providing an approach to measure and potentially recover from these errors. Specifically, we use a Bayesian approach to study the effects of algorithmic errors on key scientific questions, including the effect of siblings on children's language experience and the association between children's production and their input. By fitting a joint model of speech behavior and algorithm behavior on real and simulated data, we show that classification errors can significantly distort estimates for both the most commonly used \gls{lena}, and a slightly more accurate open-source alternative (the Voice Type Classifier from the ACLEW system). We further show that a Bayesian calibration approach for recovering unbiased estimates of effect sizes can be effective and insightful, but does not provide a fool-proof solution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。