arXiv:2409.09340cs.SDcs.AI2024-09中稿 · INTERSPEECH 2025被引 2

用可穿戴设备采集自视角语音,提升自闭症儿童-成人对话中的说话人识别准确率。

Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling

  • 通过可穿戴传感器采集自视角语音数据,构建真实互动场景下的语音样本。
  • 基于Ego4D预训练模型,使儿童与成人说话人分类准确率显著提升。
  • 为自闭症干预评估提供新方法,适合临床行为分析与智能辅助研究者。

自闭症谱系障碍(ASD)是一种神经发育障碍,表现为社交沟通困难、重复行为和感官处理异常。评估儿童在治疗过程中行为变化的重要研究方向是通过标准协议BOSCC进行儿童与临床医生的双人互动,完成一系列预设活动。理解儿童行为的关键在于自动语音理解,特别是识别谁在说话及何时说话。现有方法多依赖旁观者视角的语音采样,缺乏对自视角语音建模的研究。本研究设计实验,利用可穿戴传感器从自视角采集BOSCC互动中的语音,并探索使用Ego4D语音样本进行预训练,以增强双人互动中儿童与成人的说话人分类性能。研究结果表明,自视角语音采集与预训练策略能有效提升说话人分类准确性,具有重要应用潜力。

原文摘要 · Abstract (English)

Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by challenges in social communication, repetitive behavior, and sensory processing. One important research area in ASD is evaluating children's behavioral changes over time during treatment. The standard protocol with this objective is BOSCC, which involves dyadic interactions between a child and clinicians performing a pre-defined set of activities. A fundamental aspect of understanding children's behavior in these interactions is automatic speech understanding, particularly identifying who speaks and when. Conventional approaches in this area heavily rely on speech samples recorded from a spectator perspective, and there is limited research on egocentric speech modeling. In this study, we design an experiment to perform speech sampling in BOSCC interviews from an egocentric perspective using wearable sensors and explore pre-training Ego4D speech samples to enhance child-adult speaker classification in dyadic interactions. Our findings highlight the potential of egocentric speech collection and pre-training to improve speaker classification accuracy.

自闭症说话人识别自视角语音可穿戴设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。