提出AI心理健康工具的负责任评估框架,强调临床、社会与公平维度。
Responsible Evaluation of AI for Mental Health
- 构建三类AI心理支持类型分类体系:评估、干预、信息整合。
- 分析135篇CL论文发现,现有评估过度依赖通用指标,忽略临床有效性和用户体验。
- 倡导多方参与评价,关注安全与公平,适合临床研究者与政策制定者参考。
尽管人工智能在心理健康领域展现出巨大潜力,但当前对其工具的评估方法仍零散且与临床实践、社会背景及用户真实体验脱节。本文主张重新思考负责任的评估——衡量什么、由谁衡量、为何衡量——提出一个融合临床合理性、社会情境与公平性的跨学科框架,为评估提供结构化基础。通过对135篇近期CL论文的分析,识别出若干共性缺陷:过度依赖不反映临床有效性、治疗适宜性或用户体验的通用指标;精神健康专业人员参与有限;对安全性和公平性关注不足。为弥补这些差距,我们提出一种AI心理支持类型的分类体系——以评估、干预和信息整合为导向,每类具有不同的风险特征与评估要求,并通过案例研究展示其应用。
原文摘要 · Abstract (English)
Although artificial intelligence (AI) shows growing promise for mental health care, current approaches to evaluating AI tools in this domain remain fragmented and poorly aligned with clinical practice, social context, and first-hand user experience. This paper argues for a rethinking of responsible evaluation -- what is measured, by whom, and for what purpose -- by introducing an interdisciplinary framework that integrates clinical soundness, social context, and equity, providing a structured basis for evaluation. Through an analysis of 135 recent *CL publications, we identify recurring limitations, including over-reliance on generic metrics that do not capture clinical validity, therapeutic appropriateness, or user experience, limited participation from mental health professionals, and insufficient attention to safety and equity. To address these gaps, we propose a taxonomy of AI mental health support types -- assessment-, intervention-, and information synthesis-oriented -- each with distinct risks and evaluative requirements, and illustrate its use through case studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。