arXiv:2602.05088cs.AI2026-02被引 5

开源工具VERA-MH可自动评估AI聊天机器人自杀风险应对能力,效果接近专业医生判断。

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

  • 用真实对话模拟不同自杀风险用户,让专家和大模型按同一标准评分。
  • 专家与大模型评分一致性达0.81,证明该工具可靠稳定。
  • 适合关注AI心理健康安全的开发者、研究者和政策制定者参考。

如今数百万人使用生成式AI聊天机器人获取心理支持。尽管前景广阔,当前最紧迫的问题是这些工具是否安全。该领域缺乏经过验证的自动化基准来评估AI聊天机器人安全性,尤其是对自杀风险用户。为此提出了「心理健康的伦理与责任AI验证」(VERA-MH)评估框架。本研究通过人类验证,检验了VERA-MH评分与临床专家判断的一致性。我们模拟了具有不同自杀风险等级和表达风格的大型语言模型(LLM)用户与通用AI聊天机器人的对话,由Spring Health的持证心理健康专家独立使用VERA-MH评分量表进行打分。同时,一个基于LLM的评估器(“裁判”)也应用相同量表对同一对话进行评分。研究考察了专家间一致性、专家共识与LLM裁判之间的一致性,以及不同裁判模型间的稳定性。专家在安全性评分上表现出强一致性(校正后组内相关系数[IRR] = 0.77),确立了可靠的临床共识基准。LLM裁判与该共识高度一致(IRR = 0.81),且在不同模型和重复评估中表现稳定。对用户-代理真实性和意图风险/表达风格的匹配度评分则结果不一。这些发现支持VERA-MH作为开源、全自动的自杀风险检测与应对评估基准。由于结果反映的是该基准早期版本,未来工作应验证更新版本,评估泛化性与鲁棒性,并将VERA-MH拓展至心理健康AI安全的更多领域。

原文摘要 · Abstract (English)

Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental health is whether these tools are safe. The field currently lacks a validated, automated benchmark for evaluating AI chatbot safety, particularly for users at risk of suicide. The Validation of Ethical and Responsible AI in Mental Health (VERA-MH) evaluation was recently proposed to address this need. This human validation study examined the alignment of VERA-MH safety ratings with expert clinician judgments. We simulated conversations between large language model (LLM)-based users spanning a range of suicide risk levels and disclosure styles and general-purpose AI chatbots. Licensed mental health clinicians from Spring Health independently rated chatbot safety using the VERA-MH scoring rubric. An LLM-based evaluator ("judge") applied the same rubric to the same conversations. We examined agreement among clinicians, between clinician consensus and the LLM judge, and across different judge LLMs. Clinicians also rated user-agent realism, suicide risk, and disclosure. Clinicians showed strong agreement in safety ratings (chance-corrected inter-rater reliability [IRR] = 0.77), establishing a reliable clinical consensus reference. The LLM judge was strongly aligned with this consensus (IRR = 0.81), and ratings were stable across judge models and repeated evaluations. Ratings of user-agent realism and fidelity to intended suicide risk and disclosure styles were mixed. These findings support the reliability of VERA-MH as an open-source, fully automated benchmark for evaluating AI chatbot suicide risk detection and response. Because these results reflect an earlier version of the benchmark, future work should validate updated versions, assess generalizability and robustness, and expand VERA-MH to additional domains of AI safety in mental health.

AI安全心理健康评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。