用儿童语音语境提升亲子对话识别准确率
Context-aware child-directed speech detection from long-form recordings
- 基于儿童语音数据微调自监督模型,效果优于通用语音模型
- 引入上下文信息使平均F1提升13.8个百分点
- 在真实端到端流程中仍优于传统规则方法,适合多语言研究
在长时录音中自动区分亲子对话与成人对话,是实现儿童语言环境规模化分析的关键。现有方法通常孤立处理话语片段,且主要在英语数据上评估。本文从三个维度弥补这些不足:首先,在包含182名儿童的多语言数据集上对六种自监督模型进行微调与评估,结果表明在以儿童为中心的录音上进行领域内预训练,显著优于在成人语音上训练的模型;其次,引入相邻语境信息可显著提升分类性能,平均F1得分绝对提升13.8%;第三,将模型置于从成人语音检测到对话对象分类的完整端到端流程中评估,尽管自动分段导致性能下降,但整体仍持续优于规则基线。
原文摘要 · Abstract (English)
Automatically distinguishing child-directed speech from adult-directed speech in long-form recordings is key to scalable analyses of children's language environments. Existing approaches process utterances in isolation and have been evaluated primarily on English. We address these gaps along three dimensions. First, we fine-tune and evaluate six-self supervised models on a multilingual dataset of 182 children, showing that in-domain pre-training on child-centered recordings substantially outperforms models trained on adult speech. Second, we demonstrate that incorporating surrounding context substantially improves classification, with an absolute gain of 13.8% in average F1-score. Third, we evaluate our model in a realistic end-to-end pipeline, from adult speech detection to addressee classification, showing that performance drops under automatic segmentation but still consistently outperforms a rule-based baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。