高帧率视频中15帧采样率可使自闭症行为识别准确率达98.75%
Evaluating the Effect of Frame Rate in Sequence-Based Classification of Autism-Related Self-Stimulatory Hand Idiosyncrasies
- 用LSTM和GRU模型分析不同帧率下的行为特征
- 15帧间隔采样时GRU模型准确率达98.75%
- 水平翻转与上采样是关键增强策略,适合小样本临床应用
自闭症谱系障碍(ASD)影响全球超7500万人,但远程行为筛查的可扩展计算方法仍有限。本研究解决视频中自动检测自闭症相关自我刺激行为的两个互补挑战:(1) 确定最优序列神经网络架构与时间采样率;(2) 建模小样本行为数据集上的数据增强策略。在自刺激行为诊断(SSBD)数据集上,对姿态特征采用1、5、15、30、45、90帧间隔采样,训练LSTM与门控循环单元(GRU)模型。两种架构均超越先前卷积神经网络(CNN)基线(62–76%准确率),峰值准确率分别为97.5%(LSTM)和98.75%(GRU),在每15帧采样时达到最优。第二项研究中,在I3D迁移学习框架上测试十种数据增强策略,消融实验量化各技术边际贡献:水平翻转单独使用时准确率最高(48.78%),而去除上采样导致性能下降最大,表明其对复杂行为视频增强至关重要。个性化机器学习方法——对每个受试者在时间分割片段上训练并测试个体模型——产生一致预测(平均损失1.84,标准差0.79)。结果为数据稀缺临床场景中的视频行为分类提供具体指导。
原文摘要 · Abstract (English)
Autism spectrum disorder (ASD) affects over 75 million individuals worldwide, yet scalable computational methods for remote behavioral screening remain limited. This study addresses two complementary challenges in automated detection of autism-related self-stimulatory behaviors from video: (1) identifying the optimal sequence-based neural network architecture and temporal sampling rate, and (2) characterizing data augmentation strategies for training on small behavioral datasets. For the first objective, long short-term memory (LSTM) and gated recurrent unit (GRU) models were trained on pose-derived features from the Self-Stimulatory Behavior Diagnosis (SSBD) dataset at frame sampling intervals of 1, 5, 15, 30, 45, and 90 frames. Both architectures exceeded prior convolutional neural network (CNN) baselines (62-76% accuracy), with peak accuracies of 97.5% (LSTM) and 98.75% (GRU) at a sampling interval of every 15 frames. For the second objective, ten data augmentation strategies were applied to an I3D transfer learning pipeline, with an ablation study quantifying the marginal contribution of each technique. Horizontal flip achieved the highest standalone accuracy (48.78%), while exclusion of upsampling from the augmentation pipeline produced the largest performance degradation, indicating its necessity for complex behavioral video augmentation. A personalized machine learning approach, in which per-subject models were trained and tested on temporally split segments of each video, produced consistent predictions (mean loss 1.84, SD 0.79). These results provide practitioners with concrete guidance on architecture selection, sampling rate, and augmentation strategy for video-based behavioral classification in data-scarce clinical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。