arXiv:2607.28476cs.CL2026-07

针对西语心理疾病早期筛查难题,构建专用模型与自动标注方法。

Improving Mental Health Screening and Early Risk Detection in Spanish

论文配图:Improving Mental Health Screening and Early Risk Detection in Spanish
图 1 · 摘自论文原文
  • 用领域预训练打造三个西语心理健康专用基础模型
  • 提出增量上下文扩展法,提前识别风险信号,缩短检测延迟
  • 公开模型与数据,适合医疗AI与社会媒体分析研究者

心理疾病早期发现常受限于西语专业资源不足及社交媒体长时文本分析困难。本文提出三项贡献:首先,通过领域特定预训练构建三个西语心理健康专用基础模型;其次,提出增量上下文扩展(ICE)方法,自动识别累积信息中足以提示疾病的临界点,生成更具信息量的训练样本;第三,利用ICE生成样本微调模型,用于早期风险检测。在三个西语基准测试中,结合专用模型与ICE的方法显著优于现有技术,降低检测延迟的同时保持高性能。所有模型均已公开。

原文摘要 · Abstract (English)

Early detection of mental health disorders is often limited by the lack of specialized resources in Spanish and the difficulty of analyzing long histories of social media posts. This paper addresses these challenges through three main contributions. First, we introduce three Spanish foundational models specifically adapted to the mental health domain through domain-specific pre-training. Second, we propose Incremental Context Expansion (ICE), an automatic relabeling methodology designed for early detection. ICE identifies the point at which cumulative messages provide enough evidence of a disorder, generating more informative training samples. Third, we provide a set of fine-tuned models using the samples generated with the ICE methodology for early risk detection tasks. Our results on three Spanish benchmarks show that combining these specialized models with ICE improves the state-of-the-art, reducing detection latency while maintaining high performance. All models are publicly available.

心理健康西语NLP早期预警模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。