为心理健康聊天机器人设计安全评估框架,重点检测自杀倾向应对能力。
VERA-MH: Validation of Ethical and Responsible AI in Mental Health

- 用临床指导的角色扮演模拟危机用户对话,覆盖多重风险因素。
- 通过医生制定的流程化评分体系,评估模型在每轮对话中的响应质量。
- 可应用于主流大模型安全评测,适合关注AI伦理与医疗应用的研究者。
聊天机器人使用率上升,已进入其原本未设计的心理健康支持领域。为此,我们提出心理健康领域人工智能伦理与责任验证框架(VERA-MH),首次聚焦自杀意念(SI)风险,评估聊天机器人对潜在危机用户的响应能力。VERA-MH包含三步:对话模拟、对话评判和模型评分。首先,由另一聊天机器人基于临床指导构建的用户角色进行模拟对话,涵盖多种风险因素、人口统计特征及披露行为。评判阶段采用第二个支持模型作为大语言模型评审员,并结合临床制定的流程化评分标准。该标准以单个是/否问题逐层推进,提升判断一致性并揭示模型失效模式。最后,汇总各轮对话结果,给出聊天机器人的最终评估。本文还展示了对四大主流大模型提供商的评估结果。
原文摘要 · Abstract (English)
Chatbot usage has increased, including in fields for which they were never developed for--notably mental health support. To that end, we introduce Validations of Ethical and Responsible AI in Mental Health (VERA-MH), a novel clinically-validated evaluation for safety of chatbots in the context of mental health support. The first iteration of VERA-MH focuses on Suicidal Ideation (SI) risks, by assessing how well chatbots can responds to users that might be in crisis. VERA-MH is comprised of three steps: conversation simulation, conversation judging and model rating. First, to simulate conversations with the chatbot under evaluation, another chatbot is tasked with role-playing users based on specific personas. Such user personas have been developed under clinical guidance, to make sure that, among others, multiple risk factors, demographic characteristics and disclosure factors were represented. In the judging step, a second support model is used as an LLM-as-a-Judge, together with a clinically-developed rubric. The rubric is structured as a flow, with a single Yes/No question asked each time, to improve answers' consistency and highlight models' failure modes. In the last stage, results of each conversation are aggregated to present the final evaluation of the chatbot. Together with the framework, we present the result of the evaluations for four leading LLM providers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。