对比三种模型在活动记录图像上的精神健康分类表现,发现混合架构最稳定可靠。
Comparative Analysis of Vision Transformer, Convolutional, and Hybrid Architectures for Mental Health Classification Using Actigraphy-Derived Images
- 用三种图像模型处理腕戴活动数据生成的图像,评估其分类效果。
- CoAtNet-Tiny平均准确率最高,对抑郁和精神分裂症类别的召回率也最强。
- 适合关注精神健康诊断中模型稳定性与小样本类别表现的研究者。
本研究比较了VGG16、ViT-B/16和CoAtNet-Tiny三种基于图像的方法在利用每日活动记录识别抑郁症、精神分裂症和健康对照中的表现。来自Psyke和Depresjon数据集的腕戴活动信号被转化为30×48像素的图像,并采用三重受试者交叉验证进行评估。尽管所有模型在训练数据上均拟合良好,但在未见数据上的表现差异显著。VGG16虽稳步提升但常停在较低准确率;ViT-B/16在部分运行中表现强劲,但跨折叠结果波动明显;CoAtNet-Tiny则最为可靠,展现出最高平均准确率和最稳定的性能曲线,尤其在抑郁与精神分裂症等少数类上具备最强的精确率、召回率与F1分数。总体表明,混合架构在基于活动记录图像的精神健康分析中具有更优的一致性。
原文摘要 · Abstract (English)
This work examines how three different image-based methods, VGG16, ViT-B/16, and CoAtNet-Tiny, perform in identifying depression, schizophrenia, and healthy controls using daily actigraphy records. Wrist-worn activity signals from the Psykose and Depresjon datasets were converted into 30 by 48 images and evaluated through a three-fold subject-wise split. Although all methods fitted the training data well, their behaviour on unseen data differed. VGG16 improved steadily but often settled at lower accuracy. ViT-B/16 reached strong results in some runs, but its performance shifted noticeably from fold to fold. CoAtNet-Tiny stood out as the most reliable, recording the highest average accuracy and the most stable curves across folds. It also produced the strongest precision, recall, and F1-scores, particularly for the underrepresented depression and schizophrenia classes. Overall, the findings indicate that CoAtNet-Tiny performed most consistently on the actigraphy images, while VGG16 and ViT-B/16 yielded mixed results. These observations suggest that certain hybrid designs may be especially suited for mental-health work that relies on actigraphy-derived images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。