用多模态行为分析提升自闭症筛查准确率
Improving Autism Detection with Multimodal Behavioral Analysis
- 融合面部表情、语音语调、头部动作等多模态数据
- gaze 分类准确率从 64% 提升至 69%
- 适合需要快速筛查的临床与研究场景
由于自闭症谱系障碍(ASC)诊断复杂且资源密集,已有多种计算机辅助诊断方法通过分析患者视频中的行为线索来检测自闭症。尽管这些模型在某些数据集上表现良好,但普遍存在眼动特征表现差和现实泛化能力不足的问题。为此,我们分析了一个标准化视频数据集,包含168名自闭症患者(46%女性)和157名非自闭症参与者(46%女性),该数据集据我们所知是目前最大且最平衡的。我们对面部表情、语音语调、头部运动、心率变异性(HRV)和注视行为进行了多模态分析。为解决以往注视模型的局限性,我们引入了新的统计描述符,量化眼球注视角度的变化性,使基于注视的分类准确率从64%提升至69%,并与临床研究中关于自闭症注视回避的发现一致。采用后融合策略,最终分类准确率达到74%,证明了跨模态行为标记整合的有效性。研究结果表明,可扩展的视频基筛查工具具有支持自闭症评估的潜力。
原文摘要 · Abstract (English)
Due to the complex and resource-intensive nature of diagnosing Autism Spectrum Condition (ASC), several computer-aided diagnostic support methods have been proposed to detect autism by analyzing behavioral cues in patient video data. While these models show promising results on some datasets, they struggle with poor gaze feature performance and lack of real-world generalizability. To tackle these challenges, we analyze a standardized video dataset comprising 168 participants with ASC (46% female) and 157 non-autistic participants (46% female), making it, to our knowledge, the largest and most balanced dataset available. We conduct a multimodal analysis of facial expressions, voice prosody, head motion, heart rate variability (HRV), and gaze behavior. To address the limitations of prior gaze models, we introduce novel statistical descriptors that quantify variability in eye gaze angles, improving gaze-based classification accuracy from 64% to 69% and aligning computational findings with clinical research on gaze aversion in ASC. Using late fusion, we achieve a classification accuracy of 74%, demonstrating the effectiveness of integrating behavioral markers across multiple modalities. Our findings highlight the potential for scalable, video-based screening tools to support autism assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。