arXiv:2604.17359cs.CYcs.AI2026-04

大模型生成的心理患者个体看似合理,但整体群体与真实数据严重不符。

Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations

  • 用临床和自述两种方式让模型生成患者,逐个验证症状符合诊断标准。
  • 个体表现达标率97.3%,但群体平均抑郁分高出2.8–5.5点,治疗阈值占比翻倍。
  • 症状分布随群体变化重组,修正边际数据无法还原真实细胞层级差异。

我们让GPT-4o-mini、Gemini-3-Flash、DeepSeek-V3和GLM-4.7分别生成120个人口学群体的患者描述,共28,800条响应,采用基于NHANES微数据的调查加权PHQ-8基准进行评估。单个案例层面:97.3%的高风险表现满足DSM-5准入规则,错误率2.68%(随机预期为10.4%)。群体层面则四重失效:所有可比群体的PHQ-8均值虚高2.8至5.5分;模拟患者中18.2%达到治疗阈值(真实为7.5%);黑-白、西-白群体差异被削弱或反转;症状协方差结构随群体重构,校准无法修复;人口偏移不叠加,边际修正误差达约0.4分。此外,同一提示下,三分之一患者在两次采样间改变严重程度等级,五分之一跨越观察到治疗的临界线。四个模型共同暴露‘一致性-真实性分离’问题:个体似真,群体失真。性别认同影响症状结构与稳定性,且缺乏联邦级基准验证。患者看似合理,却不代表真实人群。

原文摘要 · Abstract (English)

Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no real one. We gave GPT-4o-mini, Gemini-3-Flash, DeepSeek-V3 and GLM-4.7 each of 120 demographic cohorts under two framings, one written as a clinician enters a patient and one as a person describes themselves, and scored all 28,800 responses against survey-weighted PHQ-8 anchors derived from NHANES microdata. Case by case the output holds up: 97.3% of elevated presentations satisfy the DSM-5 gateway rule, violating it at 2.68% against a chance null of 10.4%. As populations, four things fail at once. Every benchmarkable group returns inflated by 2.8 to 5.5 PHQ-8 points, and 18.2% of simulated patients screen at the treatment threshold against 7.5% of adults. Population Black-White and Hispanic-White disparities do not survive the simulation, with two models attenuating each gap and two flattening or inverting it. Symptom covariance reorganizes by cohort, putting the error beyond any recalibration, and demographic offsets do not stack, so a correction fitted on marginals misses the cells by about 0.4 points either way. And the answer does not hold still: at the decoding a deployment inherits, a third of patients change severity category between two draws of one prompt and one in five crosses the line from watchful waiting to treatment. We read the four together as one failure, and name the gap between case-level plausibility and population-level failure the coherence-fidelity dissociation. Individual cases clear a formal rule a case review would apply; the populations they compose fail every comparison we can construct against a real one. Gender identity carries an extreme on both symptom structure and regeneration stability, and no federal benchmark exists to check any of it. The patients look right. They do not represent real populations.

心理健康大模型审计群体偏差真实性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。