arXiv:2604.22063cs.LGcs.AI2026-04

测试大模型在精神科风险评分中的可靠性,发现无关信息会显著影响结果稳定性。

Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores

论文配图:Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores
图 1 · 摘自论文原文
  • 通过构造含无关特征的虚拟患者数据,评估提示词和噪声输入对预测的影响。
  • 加入无关临床特征后,所有模型的风险评分均显著升高且波动加剧。
  • 提示词设计会独立影响模型稳定性,提醒临床部署前需系统审计

大型语言模型(LLMs)在临床推理与风险评估中应用日益广泛,但在精神科等模糊领域其解释可靠性尚不明确。已有研究指出算法偏见与提示敏感性问题,但缺乏针对精神科场景的系统评估方法。本文提出一种下游任务可靠性审计方法,聚焦提示设计及医学无关输入对住院风险评分的影响。我们生成50个合成患者样本,每个包含15个临床相关特征及最多50个医学无关特征,结合四种提示重构方式(中性、逻辑、人文影响、临床判断)。审计了四个模型(Gemini 2.5 Flash、LLaMa 3.3 70b、Claude Sonnet 4.6、GPT-4o mini)。结果显示,引入医学无关变量导致所有模型和提示下的绝对平均风险评分显著上升,输出变异程度增加,表明预测稳定性随上下文噪声增强而下降。无关特征在多数模型-提示组合中引发不稳定性,提示词变化亦以模型依赖方式独立影响不稳定趋势。该研究量化了精神科风险评估中对非临床信息的敏感性,强调临床部署前必须开展归因稳定性和不确定性行为的系统评估。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly utilized in clinical reasoning and risk assessment. However, their interpretive reliability in critical and indeterminate domains such as psychiatry remains unclear. Prior work has identified algorithmic biases and prompt sensitivity in these systems, raising concerns about how contextual information may influence model outputs, but there remains no systematic way to assess these, especially in the psychiatric domain. We propose an approach for reliability auditing downstream LLM tasks by structuring evaluation around the impact of prompt design and the inclusion of medically insignificant inputs on predicted hospitalization risk scores, which is often the first downstream AI clinical-decision-making task. In our audit, a cohort of synthetic patient profiles (n = 50) is generated, each consisting of 15 clinically relevant features and up to 50 clinically insignificant features, across four prompt reframings (neutral, logical, human impact, clinical judgment). We audit four LLMs (Gemini 2.5 Flash, LLaMa 3.3 70b, Claude Sonnet 4.6, GPT-4o mini), and our results show that including medically insignificant variables resulted in a statistically significant increase in the absolute mean predicted hospitalization risk and output variability across all models and prompts, indicating reduced predictive stability as contextual noise increased. Clinically insignificant features had an effect on instability across many model-prompt conditions, and prompt variations independently affected the trajectory of instability in a model-dependent manner. These findings quantify how LLM-based psychiatric risk assessments are sensitive to non-clinical information, highlighting the need for systematic evaluations of attributional stability and uncertainty behavior like this before clinical deployments.

精神科AI大模型可靠性风险评估提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。