arXiv:2504.03695eess.SPcs.AI2025-04被引 3

测试可穿戴设备焦虑模型跨活动和人群的泛化能力,发现性能稳定但敏感度仍有提升空间。

Are Anxiety Detection Models Generalizable? A Cross-Activity and Cross-Population Study Using Wearables

  • 在三类焦虑场景下训练3348个模型,评估跨活动与跨人群表现
  • 最佳模型AUROC达0.62–0.73,焦虑状态召回率35.19%–74.3%
  • 模型对不同活动和人群表现较稳定,适合关注公平性与实际部署的研究者

焦虑相关活动(如公开演讲)会引发焦虑障碍患者的显著焦虑反应。近年来研究发现,通过可穿戴设备采集的心电图(ECG)和皮肤电活动(EDA)等生理信号,结合机器学习模型可在这些情境中检测焦虑。然而,这类焦虑预测模型在不同活动和多样人群间的泛化能力仍缺乏充分探索,而这一问题对评估模型偏见及建立用户信任至关重要。为此,我们招募了111名参与者,经历三种焦虑诱发任务。利用自建数据集及两个知名公开数据集,评估了模型在参与者内部(同活动与跨活动)以及跨参与者(同活动与跨活动)的泛化表现。共训练并测试超过3348个焦虑检测模型(采用六种分类器、31种特征集与18种训练-测试配置)。结果表明,三个关键指标——AUROC、焦虑状态召回率与非焦虑状态召回率——均略高于0.5基线。最佳AUROC值为0.62至0.73,焦虑类召回率范围为35.19%至74.3%。有趣的是,模型性能(以AUROC衡量)在不同活动和人群间相对稳定,尽管焦虑类召回率存在一定程度波动。

原文摘要 · Abstract (English)

Anxiety-provoking activities, such as public speaking, can trigger heightened anxiety responses in individuals with anxiety disorders. Recent research suggests that physiological signals, including electrocardiogram (ECG) and electrodermal activity (EDA), collected via wearable devices, can be used to detect anxiety in such contexts through machine learning models. However, the generalizability of these anxiety prediction models across different activities and diverse populations remains underexplored-an essential step for assessing model bias and fostering user trust in broader applications. To address this gap, we conducted a study with 111 participants who engaged in three anxiety-provoking activities. Utilizing both our collected dataset and two well-known publicly available datasets, we evaluated the generalizability of anxiety detection models within participants (for both same-activity and cross-activity scenarios) and across participants (within-activity and cross-activity). In total, we trained and tested more than 3348 anxiety detection models (using six classifiers, 31 feature sets, and 18 train-test configurations). Our results indicate that three key metrics-AUROC, recall for anxious states, and recall for non-anxious states-were slightly above the baseline score of 0.5. The best AUROC scores ranged from 0.62 to 0.73, with recall for the anxious class spanning 35.19% to 74.3%. Interestingly, model performance (as measured by AUROC) remained relatively stable across different activities and participant groups, though recall for the anxious class did exhibit some variation.

焦虑检测可穿戴设备泛化能力生理信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。