arXiv:2505.15556cs.CL2025-05Conference of the …综述被引 6

梳理多语言社交媒体心理疾病检测现状与挑战

A Survey on Multilingual Mental Disorders Detection from Social Media Data

  • 系统整理25种语言的108个心理健康数据集
  • 指出低资源语言数据稀缺与抑郁数据主导问题
  • 适合跨语言NLP与数字心理健康研究者参考

全球心理障碍患病率上升,亟需适用于多语言环境的数字筛查方法。现有研究多集中于英语数据,忽视非英语文本中的关键心理健康信号。本文综述了社交媒体数据中多语言心理障碍检测的研究进展,整理出涵盖25种语言的108个可用数据集,讨论文化差异对线上语言模式和自我披露行为的影响及其对NLP工具性能的干扰。文中指出现有挑战:低-中资源语言资源匮乏,且数据以抑郁症为主,其他障碍研究不足。为此,呼吁跨学科合作与多语言基准建设,推动全球心理健康筛查能力提升。

原文摘要 · Abstract (English)

The increasing prevalence of mental disorders globally highlights the urgent need for effective digital screening methods that can be used in multilingual contexts. Most existing studies, however, focus on English data, overlooking critical mental health signals that may be present in non-English texts. To address this gap, we present a survey of the detection of mental disorders using social media data beyond the English language. We compile a comprehensive list of 108 datasets spanning 25 languages that can be used for developing NLP models for mental health screening. In addition, we discuss the cultural nuances that influence online language patterns and self-disclosure behaviors, and how these factors can impact the performance of NLP tools. Our survey highlights major challenges, including the scarcity of resources for low- and mid-resource languages and the dominance of depression-focused data over other disorders. By identifying these gaps, we advocate for interdisciplinary collaborations and the development of multilingual benchmarks to enhance mental health screening worldwide.

心理检测多语言社交媒体NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。