为聊天式AI心理治疗师设计风险评估框架,提升安全性和可靠性。
A Risk Ontology for Evaluating AI-Powered Psychotherapy Virtual Agents
- 构建针对AI心理治疗的系统性风险分类体系
- 可识别用户认知行为变化及潜在严重后果
- 适合临床评估、安全测试与模型对比研究
大型语言模型和智能虚拟治疗师的普及为扩大心理健康服务可及性带来机遇,但也因缺乏标准化评估方法,导致用户伤害甚至自杀等严重不良事件频发。现有评估手段难以捕捉治疗过程中患者认知与行为的细微变化,这些变化可能引发后续病情恶化。本文提出一种专为对话式AI心理治疗师设计的风险本体,基于心理治疗风险文献、临床与法律专家访谈,并与DSM-5及现有评估工具(如NEQ、UE-ATR)对齐,旨在系统识别与评估用户/患者风险。提供该本体的高层概述及其理论基础,并详细讨论四大应用场景:真实用户交互监控、模拟患者评估、模型基准对比分析及异常结果发现。该框架为实现更安全、负责任的AI心理健康支持创新奠定基础。
原文摘要 · Abstract (English)
The proliferation of Large Language Models (LLMs) and Intelligent Virtual Agents acting as psychotherapists presents significant opportunities for expanding mental healthcare access. However, their deployment has also been linked to serious adverse outcomes, including user harm and suicide, facilitated by a lack of standardized evaluation methodologies capable of capturing the nuanced risks of therapeutic interaction. Current evaluation techniques lack the sensitivity to detect subtle changes in patient cognition and behavior during therapy sessions that may lead to subsequent decompensation. We introduce a novel risk ontology specifically designed for the systematic evaluation of conversational AI psychotherapists. Developed through an iterative process including review of the psychotherapy risk literature, qualitative interviews with clinical and legal experts, and alignment with established clinical criteria (e.g., DSM-5) and existing assessment tools (e.g., NEQ, UE-ATR), the ontology aims to provide a structured approach to identifying and assessing user/patient harms. We provide a high-level overview of this ontology, detailing its grounding, and discuss potential use cases. We discuss four use cases in detail: monitoring real user interactions, evaluation with simulated patients, benchmarking and comparative analysis, and identifying unexpected outcomes. The proposed ontology offers a foundational step towards establishing safer and more responsible innovation in the domain of AI-driven mental health support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。