arXiv:2601.11758cs.CLcs.AI2026-01

用可解释的语言特征,从社交媒体早期文本中精准识别焦虑倾向。

Early Linguistic Pattern of Anxiety from Social Media Using Interpretable Linguistic Features: A Multi-Faceted Validation Study with Author-Disjoint Evaluation

  • 基于语言学特征构建透明分类模型,避免黑箱决策。
  • 仅用少量发帖记录即显著优于随机分类,跨领域验证一致。
  • 适合心理筛查研究者与需要可解释AI的临床应用者。

焦虑影响全球数亿人,但大规模筛查仍受限。社交媒体语言为可扩展检测提供可能,但现有模型常缺乏可解释性、关键词鲁棒性验证和严格的用户级数据完整性保障。本研究提出一种基于语言学可解释特征的透明焦虑检测方法,并通过跨领域验证进行多维度评估。基于大量Reddit帖子数据,我们在精心筛选的子版块中划分训练、验证与测试集,训练逻辑回归分类器。评估包含特征消融、关键词屏蔽实验、焦虑组与对照组在不同密度下的差异分析,以及使用经临床访谈确诊的焦虑障碍患者进行外部验证。模型在去除情感信息或屏蔽关键词后仍保持高精度,仅凭少量发帖历史即显著优于随机分类,跨域分析结果与临床访谈数据高度一致。结果表明,透明语言学特征可支持可靠、可泛化且关键词鲁棒的焦虑检测。该框架为不同在线场景中的可解释心理健康筛查提供了可复现基线。

原文摘要 · Abstract (English)

Anxiety affects hundreds of millions of individuals globally, yet large-scale screening remains limited. Social media language provides an opportunity for scalable detection, but current models often lack interpretability, keyword-robustness validation, and rigorous user-level data integrity. This work presents a transparent approach to social media-based anxiety detection through linguistically interpretable feature-grounded modeling and cross-domain validation. Using a substantial dataset of Reddit posts, we trained a logistic regression classifier on carefully curated subreddits for training, validation, and test splits. Comprehensive evaluation included feature ablation, keyword masking experiments, and varying-density difference analyses comparing anxious and control groups, along with external validation using clinically interviewed participants with diagnosed anxiety disorders. The model achieved strong performance while maintaining high accuracy even after sentiment removal or keyword masking. Early detection using minimal post history significantly outperformed random classification, and cross-domain analysis demonstrated strong consistency with clinical interview data. Results indicate that transparent linguistic features can support reliable, generalizable, and keyword-robust anxiety detection. The proposed framework provides a reproducible baseline for interpretable mental health screening across diverse online contexts.

焦虑检测可解释性社交媒体语言特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。