用多模态可穿戴设备精准识别拔发、抠皮等强迫行为。
Deep Multimodal Wearable Sensor Fusion for Detection of Body-Focused Repetitive Behaviors

- 融合加速度、温度和距离传感器数据,用深度学习模型捕捉动作特征。
- 二分类F1达0.985,九类行为识别平均F1为0.700,显著优于单模态方法。
- 适合心理健康监测与临床干预研究,可实现连续无感评估。
拔发、抠皮等身体聚焦重复行为是强迫症与焦虑症的常见强迫性动作,其早期客观检测困难,因动作细微且易与正常动作重叠。本文基于儿童心智研究所使用的Helios可穿戴设备采集的腕部传感器数据,融合惯性测量单元、热电堆传感器与飞行时间传感器,捕捉运动学、热信号与空间接近度信息,构建多模态深度学习框架。该框架结合卷积神经网络与门控循环单元,辅以模态专用自编码器与后融合分类器,有效建模时空动态。在二分类任务中,检测行为与非目标活动的F1得分为0.985,受试者工作特征曲线下面积(AUC)达0.997;在九类区分任务中,宏平均F1为0.700,AUC为0.963,优于单模态基线。事后可解释性分析(Shapley值)显示,飞行时间与惯性模态主导判别能力,分别捕捉空间接近度与动态运动特征;层次聚类表明误分类主要由动作部位解剖位置决定。结果表明,多模态传感融合可实现高精度、客观、连续的行为监测,为实时可穿戴心理健康诊断与个性化干预提供基础。
原文摘要 · Abstract (English)
Body-focused repetitive behaviors, such as hair pulling and skin picking, are compulsive motor actions commonly associated with obsessive-compulsive and anxiety disorders. Their early, objective detection remains difficult because the movements are subtle and overlap with ordinary, non-pathological gestures. We developed and evaluated a multimodal deep learning framework to detect and classify these behaviors from wrist-worn sensor data. The data, collected by the Child Mind Institute using the Helios wrist-worn device, combine inertial measurement units, thermopile sensors, and time-of-flight sensors, capturing kinematic, thermal, and proximity information. The framework combined a convolutional neural network with a gated recurrent unit, alongside modality-specific autoencoders and a late-fusion classifier, to exploit temporal and spatial dynamics. It achieved an F1 score of 0.985 and an area under the receiver operating characteristic curve of 0.997 for binary detection, distinguishing these behaviors from other activities, and a macro-averaged F1 score of 0.700 with an area under the curve of 0.963 across a nine-class scheme that distinguished each individual behavior from a single grouped Non-Target class, improving over single-modality baselines. Post-hoc interpretability based on Shapley additive explanations showed that the time-of-flight and inertial modalities dominated discriminative power by capturing spatial proximity and dynamic movement, while hierarchical clustering indicated that misclassifications were driven primarily by the anatomical region of the gesture. These findings demonstrate that multimodal sensor fusion enables accurate, objective, and continuous behavioral monitoring. This work establishes a foundation for real-time, wearable-assisted mental health diagnostics and personalized interventions in biomedical research and clinical care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。