通过四要素框架系统测试大模型心理健康回复中的幻觉与遗漏风险
Disentangling Prompt Element Level Risk Factors for Hallucinations and Omissions in Mental Health LLM Responses

- 构建用户、话题、上下文、语气四要素可控的提示框架,实现压力情境下的系统化测试
- 13.2%回复存在关键安全信息遗漏,危机和自杀意念类问题尤为严重
- 上下文和语气是导致错误的核心因素,适合安全评估与模型优化研究者参考
心理健康问题常在临床场景外表达,尤其在高压力求助情境中,可能需要关键安全指导。消费健康信息系统越来越多地采用大语言模型(LLMs)进行心理健康问答,但现有评估往往忽略叙述性、高压力询问。本文提出UTCO(用户、话题、上下文、语气)提示构建框架,将询问表示为四个可调控元素,用于系统性压力测试。基于2,075个UTCO生成的提示,评估了Llama 3.3模型,标注出幻觉(虚构或错误的临床内容)和遗漏(缺失临床上必要或安全关键的指导)。结果显示,6.5%的回复存在幻觉,13.2%存在遗漏,且遗漏主要集中在危机和自杀意念类提示中。通过回归分析、元素特异性匹配及相似度匹配比较,失败最一致地与上下文和语气相关,而用户背景特征在平衡后未显示出系统性差异。这些发现支持将遗漏作为首要安全评估指标,并推动超越静态基准题集的评估范式。
原文摘要 · Abstract (English)
Mental health concerns are often expressed outside clinical settings, including in high-distress help seeking, where safety-critical guidance may be needed. Consumer health informatics systems increasingly incorporate large language models (LLMs) for mental health question answering, yet many evaluations underrepresent narrative, high-distress inquiries. We introduce UTCO (User, Topic, Context, Tone), a prompt construction framework that represents an inquiry as four controllable elements for systematic stress testing. Using 2,075 UTCO-generated prompts, we evaluated Llama 3.3 and annotated hallucinations (fabricated or incorrect clinical content) and omissions (missing clinically necessary or safety-critical guidance). Hallucinations occurred in 6.5% of responses and omissions in 13.2%, with omissions concentrated in crisis and suicidal ideation prompts. Across regression, element-specific matching, and similarity-matched comparisons, failures were most consistently associated with context and tone, while user-background indicators showed no systematic differences after balancing. These findings support evaluating omissions as a primary safety outcome and moving beyond static benchmark question sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。