用多跳问答发现大模型在心理健康中的偏见放大与沉默
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
- 设计多跳问答框架,系统检测心理医疗大模型的交叉性偏见
- 四款模型均存在情感、人口统计与病情上的系统性偏差
- 提出角色扮演和显式去偏方法,减少66%-94%偏见
大型语言模型(LLMs)在心理健康领域可能传播加剧污名化并伤害边缘群体的偏见。尽管已有研究指出相关问题,但系统性检测交叉性偏见的方法仍有限。本文提出多跳问答(MHQA)框架,分析来自可解释心理健康指令(IMHI)数据集的内容,涵盖症状表现、应对机制与治疗方式。通过年龄、种族、性别与社会经济地位的系统标注,探究人口交叉维度下的偏见模式。评估了Claude 3.5 Sonnet、Jamba 1.6、Gemma 3与Llama 4四款模型,发现其在情感倾向、人口特征与精神状况方面存在系统性差异。所提MHQA方法优于传统手段,能识别偏见在连续推理中被放大的关键节点。采用角色扮演模拟与显式去偏技术,结合BBQ数据集少样本提示,实现66%-94%的偏见降低。研究揭示了大模型再现心理健康偏见的关键环节,为公平人工智能发展提供可操作洞见。
原文摘要 · Abstract (English)
Large Language Models (LLMs) in mental healthcare risk propagating biases that reinforce stigma and harm marginalized groups. While previous research identified concerning trends, systematic methods for detecting intersectional biases remain limited. This work introduces a multi-hop question answering (MHQA) framework to explore LLM response biases in mental health discourse. We analyze content from the Interpretable Mental Health Instruction (IMHI) dataset across symptom presentation, coping mechanisms, and treatment approaches. Using systematic tagging across age, race, gender, and socioeconomic status, we investigate bias patterns at demographic intersections. We evaluate four LLMs: Claude 3.5 Sonnet, Jamba 1.6, Gemma 3, and Llama 4, revealing systematic disparities across sentiment, demographics, and mental health conditions. Our MHQA approach demonstrates superior detection compared to conventional methods, identifying amplification points where biases magnify through sequential reasoning. We implement two debiasing techniques: Roleplay Simulation and Explicit Bias Reduction, achieving 66-94% bias reductions through few-shot prompting with BBQ dataset examples. These findings highlight critical areas where LLMs reproduce mental healthcare biases, providing actionable insights for equitable AI development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。