arXiv:2608.26138cs.CLcs.LG2026-08

跨平台心理健康文本模型表现严重下滑,五维公平性审计揭示系统性缺陷。

Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media

论文配图:Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media
图 1 · 摘自论文原文
  • 构建五轴公平性评估框架,覆盖性能、校准、显著性等维度
  • 跨平台AUC下降超30%,校准误差飙升至0.54,预测公平性失衡
  • 适合关注心理健康AI伦理与跨平台部署的研究者和开发者

本文提出跨平台公平性评估(CPFE)框架,涵盖判别性能、校准性、统计显著性、预测公平性和归因稳定性五个维度。将四种Transformer模型(BERT、RoBERTa、Emotion-DistilRoBERTa、GoEmotions-RoBERTa)在Kaggle心理健康的语料库(n=35,556)上训练,并在Reddit(n=6,257)和Twitter(n=2,883)测试集上评估,情绪标签映射为临床代理。所有三组独立验证的模型均显示显著的跨平台AUC下降:在Reddit上下降30.3-35.4%,在Twitter上下降37.9-39.5%,相较域内性能(AUC 0.983-0.987)。校准失败严重,ECE从域内0.056-0.060升至Reddit的0.196-0.229和Twitter的0.499-0.542。平台特定温度缩放可降低平均ECE达88.0%而不影响判别性能(|delta AUC|<0.01),表明故障模式可分离。预测公平性分析显示显著跨平台差异(原始DI < 0.17;先验偏移调整后为0.11-0.29),在Reddit上心理健康代理类别的等几率差异为0.753-0.830,在Twitter上焦虑类为0.755-0.831。归因稳定性分析显示,跨平台词汇重叠几乎为零(16对中14对Jaccard J=0,K=10)。这些发现支持在异构环境中对心理健康NLP系统进行五轴全面验证。单种子微调实验显示,目标平台标签作为训练信号比校准信号更有效,平均AUC提升0.216。

原文摘要 · Abstract (English)

We introduce the Cross-Platform Fairness Evaluation (CPFE) framework -- a five-axis audit protocol covering discriminative performance, calibration, statistical significance, prediction equity, and attribution stability -- and apply it to four transformer models (BERT, RoBERTa, Emotion-DistilRoBERTa, GoEmotions-RoBERTa) trained on a Kaggle mental health corpus (n=35,556) and evaluated on Reddit (n=6,257) and Twitter (n=2,883) test sets with emotion labels mapped to clinical proxies. All three independently evaluated models exhibit consistent and substantial cross-platform AUC degradation (30.3-35.4% on Reddit, 37.9-39.5% on Twitter) relative to within-platform performance (AUC 0.983-0.987), confirmed across five independent training seeds. Calibration failure is concurrent and severe: ECE rises from 0.056-0.060 in-domain to 0.196-0.229 on Reddit and 0.499-0.542 on Twitter. Platform-specific temperature scaling reduces mean ECE by 88.0% without altering discriminative performance (mean |delta AUC|<0.01), confirming separable failure modes. Prediction equity analysis reveals large cross-platform disparities (raw DI < 0.17; prior-shift-adjusted DI: 0.11-0.29 on Reddit), with equalized odds differences of 0.753-0.830 for mental health proxy classes on Reddit and 0.755-0.831 for anxiety on Twitter. Attribution stability analysis shows near-complete vocabulary divergence across platforms (Jaccard J=0 in 14/16 model-class pairs at K=10). These findings support treating cross-platform validation across all five CPFE axes as a standard requirement for mental health NLP systems in heterogeneous environments. In a single-seed fine-tuning experiment, mean AUC improved by 0.216, suggesting target-platform labels provide greater benefit as training signal than as calibration signal.

心理健康NLP公平性审计跨平台泛化模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。