构建三层次信任框架,解决心理健康AI中人、交互与AI层的信任错配问题。
Trust Stack for Mental Health AI: A Survey of Calibration across Human, Interaction, and AI Layers

- 提出人类-交互-AI三层信任模型,分层解析不同主体的信赖机制。
- 分析61篇论文发现:情感化聊天机器人引发高信任但安全不足,安全系统因透明度低被低估。
- 适合关注心理健康AI伦理、可解释性及跨领域协作的研究者和开发者。
基于语言的AI正广泛用于心理健康支持,但信任评估在跨学科间存在操作脱节:自然语言处理与AI研究关注鲁棒性、安全性、隐私和可解释性,而心理治疗、人机交互与监管研究则强调治疗契合度、生活体验、共情与依赖感。具共情能力的聊天机器人可能激发强烈用户信任,却缺乏相应安全保障;而更安全的系统因边界不透明而遭低估,形成信任校准缺口,且无单一领域能独揽责任。通过对61篇论文的结构化综述,本文将该领域归纳为三层框架:(L1)以人类为中心的信任,(L2)以交互为中心的可信性,(L3)以AI为中心的可信性,并映射五类利益相关方视角。文章提出一个社会技术协同的可信心理健康AI研究议程,主张核心目标应从提升感知信任转向实现人类信任与交互及AI层面可信性的精准对齐。
原文摘要 · Abstract (English)
Language-based AI is increasingly deployed for mental health support, yet trust is evaluated in interdisciplinary but operationally misaligned ways: NLP and AI work measures robustness, safety, privacy, and explanations, while psychotherapy, HCI, and regulatory work emphasize therapeutic fidelity, lived experience, empathy, and reliance. Empathetic chatbots can elicit strong user trust without commensurate safety, while safer systems are under-trusted when their boundaries are opaque, a calibration gap no single community owns. Through a structured scoping synthesis of 61 papers, we survey this landscape into a three-layer framework separating (L1) human-oriented trust, (L2) interaction-oriented trustworthiness, and (L3) AI-oriented trustworthiness, and map five stakeholder perspectives onto these layers. We outline a research agenda for building socio-technically aligned trustworthy AI for mental health support, highlighting that the central objective should shift from maximizing perceived trust to calibrating human trust to demonstrated interaction- and AI-level trustworthiness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。