LLM误判心理困扰,对特定社群的表达反应过度。
Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts

- 用社区视角标注1198条Reddit帖子,建模不同群体对情绪表达的判断差异。
- 开源LLM在无到轻度困扰场景下准确率仅31%-44%,多数误判为有困扰。
- 模型存在普遍性困扰偏见,比人类更易高估社群中的心理危机。
心理困扰的判断具有社会情境性:何为值得关注取决于社群对情感表达、脆弱性和求助行为的规范。然而,用于检测心理困扰的大语言模型(LLMs)通常遵循单一、统一的标准。本研究通过一种视角主义标注实验,让321名参与者对来自6个身份社群的1,198条Reddit帖子进行9,587次评估,生成社区特异性标签。在情境化同群条件下,评估者与本社群的认同度略高于非情境化异群评估者(OR = 1.18),该效应在不同社群间差异显著。随后,我们以这些标签评估了九种开源权重和四种前沿模型配置。结果显示,开源模型系统性高估困扰程度:当社群认为无至轻度困扰时,其准确率仅为31%-44%,主要表现为假阳性。GPT-5与Gemini 2.5 Pro即使整体样本的过/低估率混合,仍呈现无至轻度困扰的膨胀趋势;Claude Opus 4则更为保守。这一现象并非简单反映外部观察立场——非情境化异群人类平均判断几乎对称(18%高估,19%低估)。相反,那些高估无至轻度案例的模型表现出超越同群及异群人类判断的内在困扰先验。这些发现对心理健康领域中人工智能的公平部署具有重要意义,提示错误校准的困扰检测可能不均衡地影响被评估社群。
原文摘要 · Abstract (English)
Judgments about psychological distress are socially situated: what counts as concerning hinges on community norms around emotional expression, vulnerability, and help-seeking. Yet large language models (LLMs) used for distress detection are typically aligned to a single, undifferentiated standard. How well do these models capture the perspectives of the communities whose language they assess? We address this question through a perspectivist annotation study in which 321 participants provided 9,587 judgments on 1,198 Reddit posts spanning six identity-based communities, yielding community-specific labels. Raters in the contextualized in-group condition show a modest tendency to agree more with their community than uncontextualized out-group raters (OR = 1.18), an effect varying significantly across communities. We then evaluate nine open-weight LLM configurations and four frontier configurations against these labels. Open-weight LLMs systematically over-estimate distress: when communities perceive none-to-mild distress, these models achieve only 31-44% accuracy, predominantly producing false positives. GPT-5 and Gemini 2.5 Pro show the same none-to-mild inflation even when their full-sample over/under rates are mixed, while Claude Opus 4 is more conservative. This pattern does not simply mirror an outsider reading position: uncontextualized out-group human aggregates were nearly symmetric, with 18% over-estimation versus 19% under-estimation. Instead, the models that inflate none-to-mild cases exhibit a distress prior that exceeds both contextualized in-group and uncontextualized out-group human judgments. These findings have implications for equitable AI deployment in mental health contexts, where miscalibrated distress detection may unevenly affect the communities being assessed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。