用双曲空间提升少样本焦虑检测,仅靠少量数据就能准确识别
HCFSLN: Adaptive Hyperbolic Few-Shot Learning for Multimodal Anxiety Detection
- 在双曲空间中建模情绪特征,增强不同类别的区分度
- 在108人数据集上达到88%准确率,比最优基线高14%
- 适合医疗少样本场景,尤其适用于语音+生理+视频多模态分析
焦虑障碍影响全球数百万人群,传统诊断依赖临床访谈,而机器学习模型因数据有限易过拟合。大规模数据收集成本高、耗时长,限制了应用。为此,我们提出双曲曲率少样本学习网络(HCFSLN),一种用于多模态焦虑检测的新型少样本学习框架,融合语音、生理信号与视频数据。HCFSLN通过双曲嵌入、跨模态注意力和自适应门控网络增强特征可分性,实现低数据下的鲁棒分类。我们从108名参与者收集了多模态焦虑数据集,并与六种少样本学习基线进行对比,取得88%的准确率,优于最佳基线14%。结果表明双曲空间能有效建模焦虑相关的语音模式,验证了少样本学习在焦虑分类中的潜力。
原文摘要 · Abstract (English)
Anxiety disorders impact millions globally, yet traditional diagnosis relies on clinical interviews, while machine learning models struggle with overfitting due to limited data. Large-scale data collection remains costly and time-consuming, restricting accessibility. To address this, we introduce the Hyperbolic Curvature Few-Shot Learning Network (HCFSLN), a novel Few-Shot Learning (FSL) framework for multimodal anxiety detection, integrating speech, physiological signals, and video data. HCFSLN enhances feature separability through hyperbolic embeddings, cross-modal attention, and an adaptive gating network, enabling robust classification with minimal data. We collected a multimodal anxiety dataset from 108 participants and benchmarked HCFSLN against six FSL baselines, achieving 88% accuracy, outperforming the best baseline by 14%. These results highlight the effectiveness of hyperbolic space for modeling anxiety-related speech patterns and demonstrate FSL's potential for anxiety classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。