用视频分析技术自动区分自闭症儿童与正常发育儿童的重复动作。
Advanced Gesture Recognition for Autism Spectrum Disorder Detection: Integrating YOLOv7, Video Augmentation, and VideoMAE for Naturalistic Video Analysis
- 结合YOLOv7检测动作,通过数据增强和VideoMAE建模时序特征。
- 在自然场景视频上达到95%准确率,优于现有方法。
- 适合做自闭症早期筛查或临床辅助诊断的研究者使用。
深度学习与无接触传感技术显著推动了医疗健康领域的人类行为自动化评估。针对自闭症谱系障碍(ASD),重复性运动行为如旋转、撞头、甩臂是关键诊断指标。本研究聚焦于在自然非受控环境下,通过分析视频区分自闭症儿童与典型发育(TD)同龄人。基于公开的自我刺激行为数据集(SSBD),将分类任务设为二分类问题:ASD vs. TD,依据典型的重复性动作。采用融合YOLOv7检测、大量视频增强与VideoMAE框架的流水线,该方法通过高比例掩码与重建策略高效捕捉时空特征。所提方法实现95%准确率、0.93精确率、0.94召回率与0.94 F1分数,显著超越先前最优结果。结果表明,在自然环境中结合先进目标检测、鲁棒数据增强与基于掩码自编码器的视频建模,可实现可靠的ASD与TD分类。
原文摘要 · Abstract (English)
Deep learning and contactless sensing technologies have significantly advanced the automated assessment of human behaviors in healthcare. In the context of autism spectrum disorder (ASD), repetitive motor behaviors such as spinning, head banging, and arm flapping are key indicators for diagnosis. This study focuses on distinguishing between children with ASD and typically developed (TD) peers by analyzing videos captured in natural, uncontrolled environments. Using the publicly available Self-Stimulatory Behavior Dataset (SSBD), we address the classification task as a binary problem, ASD vs. TD, based on stereotypical repetitive gestures. We adopt a pipeline integrating YOLOv7-based detection, extensive video augmentations, and the VideoMAE framework, which efficiently captures both spatial and temporal features through a high-ratio masking and reconstruction strategy. Our proposed approach achieves 95% accuracy, 0.93 precision, 0.94 recall, and 0.94 F1 score, surpassing the previous state-of-the-art by a significant margin. These results demonstrate the effectiveness of combining advanced object detection, robust data augmentation, and masked autoencoder-based video modeling for reliable ASD vs. TD classification in naturalistic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。