arXiv:2604.25053cs.CLcs.AI2026-04

分析大模型推理过程,挖掘其隐藏的心理健康偏见

Analyzing LLM Reasoning to Uncover Mental Health Stigma

论文配图:Analyzing LLM Reasoning to Uncover Mental Health Stigma
图 1 · 摘自论文原文
  • 通过解析模型中间推理步骤,发现隐性偏见
  • 比传统选择题方法多发现近三倍的歧视性表述
  • 适合关注AI伦理与心理健康应用的研究者

尽管大型语言模型(LLMs)在心理健康领域应用日益广泛,但近期研究显示,它们可能对心理疾病患者表现出偏见。现有评估主要依赖多选题(MCQs),难以捕捉模型内在逻辑中的偏见。本文分析LLM的中间推理步骤,揭示其隐藏的歧视性语言及背后逻辑。我们结合临床专业知识,对针对心理疾病患者的歧视性语言进行分类,并用于标记模型推理中的问题表述。同时,我们评估这些表述的严重程度,区分明显偏见与较隐蔽的非即时伤害性偏见。为拓展推理范围,我们还扩展了现有的心理健康偏见基准,纳入更多心理状况。结果表明,评估模型推理不仅能暴露远超传统方法的偏见,还能揭示模型在心理健康理解上的逻辑缺陷。

原文摘要 · Abstract (English)

While large language models (LLMs) are increasingly being explored for mental health applications, recent studies reveal that they can exhibit stigma toward individuals with psychological conditions. Existing evaluations of this stigma primarily rely on multiple-choice questions (MCQs), which fail to capture the biases embedded within the models' underlying logic. In this paper, we analyze the intermediate reasoning steps of LLMs to uncover hidden stigmatizing language and the internal rationales driving it. We leverage clinical expertise to categorize common patterns of stigmatizing language directed at individuals with psychological conditions and use this framework to identify and tag problematic statements in LLM reasoning. Furthermore, we rate the severity of these statements, distinguishing between overt prejudice and more subtle, less immediately harmful biases. To broaden the reasoning domain and capture a wider array of patterns, we also extend an existing mental health stigma benchmark by incorporating additional psychological conditions. Our findings demonstrate that evaluating model reasoning not only exposes substantially more stigma than traditional MCQ-based methods but it helps to identify the flaws in the LLMs' logic and their understanding of mental health conditions.

大模型偏见心理健康推理分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。