arXiv:2604.25840cs.CLcs.AI2026-04被引 3

构建可解释的抑郁患者模拟评估框架,揭示现有模型行为缺陷

PSI-Bench: Towards Clinically Grounded and Interpretable Evaluation of Depression Patient Simulators

论文配图:PSI-Bench: Towards Clinically Grounded and Interpretable Evaluation of Depression Patient Simulators
图 1 · 摘自论文原文
  • 提出多层级临床诊断评估框架,覆盖对话回合、完整对话与群体层面
  • 发现模拟器响应过长、情感转变过快且多样性不足,存在统一情绪轨迹
  • 验证框架与专家判断高度一致,适合指导心理训练系统设计

患者模拟器在心理健康培训中日益重要,但模拟抑郁患者极具挑战,因安全约束和高个体差异性导致行为真实性难以保障。现有评估依赖提示不明确的LLM裁判,且缺乏对行为多样性的衡量。本文提出PSI-Bench,一个自动化的可解释评估框架,从对话回合、完整对话及群体层面,对抑郁患者模拟器行为进行临床基础诊断。基于该框架,我们对比了两个模拟框架下七种LLM的表现,发现其生成响应过长、词汇多样性高,但行为变异性低,情感快速由负面转向正面,呈现统一的情绪演变路径。人类研究进一步证实本基准与专家判断高度一致。结果表明,模拟框架的影响大于模型规模,当前模拟器存在关键缺陷。本工作为未来模拟器的设计与评估提供了可解释、可扩展的基准。

原文摘要 · Abstract (English)

Patient simulators are gaining traction in mental health training by providing scalable exposure to complex and sensitive patient interactions. Simulating depressed patients is particularly challenging, as safety constraints and high patient variability complicate simulations and underscore the need for simulators that capture diverse and realistic patient behaviors. However, existing evaluations heavily rely on LLM-judges with poorly specified prompts and do not assess behavioral diversity. We introduce PSI-Bench, an automatic evaluation framework that provides interpretable, clinically grounded diagnostics of depression patient simulator behavior across turn-, dialogue-, and population-level dimensions. Using PSI-Bench, we benchmark seven LLMs across two simulator frameworks and find that simulators produce overly long, lexically diverse responses, show reduced variability, resolve emotions too quickly, and follow a uniform negative-to-positive trajectory. We also show that the simulation framework has a larger impact on fidelity than the model scale. Results from a human study demonstrate that our benchmark is strongly aligned with expert judgments. Our work reveals key limitations of current depression patient simulators and provides an interpretable, extensible benchmark to guide future simulator design and evaluation.

心理模拟评估框架抑郁症LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。