arXiv:2506.11075eess.AScs.LG2025-06被引 4

梳理15年儿童佩戴录音设备研究,揭示数据有效性挑战与改进路径

Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity

  • 基于全天候儿童穿戴录音,减少观察者偏差
  • 指出自动化标注中误差来源及数据质量风险
  • 提供可操作的质量评估策略,适合语言研究者使用

儿童佩戴设备采集的音频记录是儿童语言研究的基础工具。全天候长时录音能以最小观察者偏差捕捉儿童的语言输入与产出,具备高有效性的潜力。然而,海量数据需依赖自动化分析提取关键指标,这带来了标注误差与结果误读的风险。本文系统总结该技术十五年来的研究成果,梳理现有资源,并揭示多种威胁自动标注准确性的误差源。为此,提出一系列可用于评估数据质量的诊断性指标,虽无法实现完全自动化质量控制,但提供了研究人员优化数据采集和合理解释分析结果的实用策略。

原文摘要 · Abstract (English)

Audio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observer bias, and therefore high validity. The sheer volume of resulting data necessitates automated analysis to extract relevant metrics for researchers and clinicians. This paper summarizes collective knowledge on this technique, providing entry points to existing resources. We also highlight various sources of error that threaten the accuracy of automated annotations and the interpretation of resulting metrics. To address this, we propose potential troubleshooting metrics to help users assess data quality. While a fully automated quality control system is not feasible, we outline practical strategies for researchers to improve data collection and contextualize their analyses.

儿童语言语音记录数据质量自动化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。