多语言抑郁检测中,用集成教师模型提升弱监督学习效果。
Uncertainty-aware Semi-supervised Ensemble Teacher Framework for Multilingual Depression Detection
- 用多个教师模型软投票生成伪标签,结合不确定性筛选
- 在阿拉伯语、孟加拉语等4种语言上超越基线,提升跨语言鲁棒性
- 适合标注数据少的场景,可用于大规模心理健康监测
从社交媒体文本中检测抑郁仍具挑战,源于语言风格差异、非正式表达及多数语言缺乏标注数据。为此,我们提出Semi-SMDNet——一种强健的半监督多语言抑郁检测网络。该框架融合教师-学生伪标签、集成学习与数据增强。采用一组教师模型,通过软投票聚合预测结果,并引入基于不确定性的阈值过滤低置信度伪标签,以降低噪声、提升学习稳定性。同时使用置信度加权训练策略,聚焦可靠伪标签样本,显著增强跨语言鲁棒性。在阿拉伯语、孟加拉语、英语和西班牙语数据集上的测试表明,该方法持续优于强基线,显著缩小资源丰富与稀缺设置间的性能差距。详细实验与分析证实框架有效,适用于多种场景,具备在标注资源有限条件下实现可扩展的跨语言心理健康监测潜力。
原文摘要 · Abstract (English)
Detecting depression from social media text is still a challenging task. This is due to different language styles, informal expression, and the lack of annotated data in many languages. To tackle these issues, we propose, Semi-SMDNet, a strong Semi-Supervised Multilingual Depression detection Network. It combines teacher-student pseudo-labelling, ensemble learning, and augmentation of data. Our framework uses a group of teacher models. Their predictions come together through soft voting. An uncertainty-based threshold filters out low-confidence pseudo-labels to reduce noise and improve learning stability. We also use a confidence-weighted training method that focuses on reliable pseudo-labelled samples. This greatly boosts robustness across languages. Tests on Arabic, Bangla, English, and Spanish datasets show that our approach consistently beats strong baselines. It significantly reduces the performance gap between settings that have plenty of resources and those that do not. Detailed experiments and studies confirm that our framework is effective and can be used in various situations. This shows that it is suitable for scalable, cross-language mental health monitoring where labelled resources are limited.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。