arXiv:2501.01594cs.CLcs.AI2025-01被引 10

构建多维度患者模拟框架,评估精神科对话机器人临床表现。

PSYCHE: A Multi-faceted Patient Simulation Framework for Evaluation of Psychiatric Assessment Conversational Agents

  • 基于多维精神健康特征模拟真实患者行为
  • 10位精神科专家验证框架有效性,结果可量化
  • 适合研究者评估精神科对话系统临床适用性

大语言模型(LLMs)的进展推动了生成类人响应的对话智能体发展。由于精神科评估依赖复杂的医患对话交互,近年来出现了旨在模拟精神科医生角色的基于LLM的精神科评估对话智能体(PACAs)。然而,目前仍缺乏标准化方法来评估PACAs与患者互动的临床适宜性。本文提出PSYCHE框架,实现对PACAs的临床相关性、伦理安全性、成本效益和量化评估。该框架通过基于多维度精神健康构念的患者模拟,定义模拟患者的个人档案、病史与行为模式,使PACAs能够对其开展评估。我们通过10位持证精神科医生参与的研究验证了该框架的有效性,并结合对模拟患者语句的深入分析进行支持。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have accelerated the development of conversational agents capable of generating human-like responses. Since psychiatric assessments typically involve complex conversational interactions between psychiatrists and patients, there is growing interest in developing LLM-based psychiatric assessment conversational agents (PACAs) that aim to simulate the role of psychiatrists in clinical evaluations. However, standardized methods for benchmarking the clinical appropriateness of PACAs' interaction with patients still remain underexplored. Here, we propose PSYCHE, a novel framework designed to enable the 1) clinically relevant, 2) ethically safe, 3) cost-efficient, and 4) quantitative evaluation of PACAs. This is achieved by simulating psychiatric patients based on a multi-faceted psychiatric construct that defines the simulated patients' profiles, histories, and behaviors, which PACAs are expected to assess. We validate the effectiveness of PSYCHE through a study with 10 board-certified psychiatrists, supported by an in-depth analysis of the simulated patient utterances.

精神科对话评估框架患者模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。