arXiv:2608.20353cs.CLcs.AI2026-08ACL

发现心理健康NLP模型在不同标注源下因词汇干扰产生性能下降

The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

  • 构建三通道诊断框架,分离词汇、语法和心理语言风格特征
  • 发现添加词汇特征会显著降低人工标注数据上的分类效果(平均降0.072)
  • 提出差异性偏差度量指标,适用于检测标注来源导致的模型偏见

计算心理健康分类器在分布偏移下表现退化,因人工标注与远程监督管道偏好不同语言信号。本文提出TSS(三通道压力探测器)多通道诊断框架,将文本分解为(A)词汇字符n-gram、(B)小型内容稀疏的形态句法通道、(C)154特征的心理语言风格通道。在四个英文数据集(N=12,906)上,TSS揭示词汇干扰效应:向风格通道加入词汇特征会导致人工标注数据上宏观F1下降(均值0.072,p<10^-4),而自动标注数据无此现象。提出度量分歧(DoD),一种源自计量经济学的差中差统计量,用于标签源审计,实例级自助推断;主估计值为DoD(BC-A) = 0.0374,95%置信区间[0.0097, 0.0651],p=0.0032。平台分层的仅推特数据DoD重现该模式:DoD-Tw(BC-A) = +0.096(p<0.001),DoD-Tw(AC-A) = -0.089(p<0.001)。干预性掩码(pos_only)在人类数据集上破坏内容词后,仍保持通道C约95-99%性能,表明风格通道不主要依赖词汇表面形式。TSS定位为诊断审计框架,非临床筛查工具:在泛化声称前识别标签源特异性捷径学习。

原文摘要 · Abstract (English)

Computational mental health (CMH) classifiers often degrade under distribution shift because human annotators and distant-supervision pipelines reward different linguistic signals. We introduce TSS (Triple-Stream Stress probe), a multi-channel diagnostic framework that decomposes text into (A) lexical character n-grams, (B) a small, mostly content-free morpho-syntactic channel, and (C) a 154-feature psycholinguistic style channel. Across four English datasets (N=12,906), TSS reveals a lexical interference effect: adding lexical features to the style channel reduces Macro-F1 on human-labeled data (mean drop 0.072, p<10^-4) but not on auto-labeled data. We propose Degree of Divergence (DoD), a difference-in-differences statistic adapted from econometrics for label-source auditing, with instance-level bootstrap inference; the headline estimate is DoD(BC-A) = 0.0374, 95% CI [0.0097, 0.0651], p=0.0032. A platform-stratified Twitter-only DoD (which removes the Reddit vs. Twitter contrast) reproduces the pattern with bootstrap inference: DoD-Tw(BC-A) = +0.096 (p<0.001) and DoD-Tw(AC-A) = -0.089 (p<0.001). Interventional masking (pos_only) retains ~95-99% of Channel C's performance after destroying content words on human datasets, indicating that the style channel does not rely primarily on lexical surface form. TSS is positioned as a diagnostic audit framework, not a clinical screening tool: it flags label-source-specific shortcut learning before generalization claims are made.

心理健康NLP标签偏差诊断框架词汇干扰

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。