arXiv:2607.07766cs.AI2026-07

为AI心理健康助手建立三层次安全对齐标准,确保长期无害且有益。

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

  • 从临床规范出发,构建价值、训练与监督三层次对齐机制。
  • 提出'对齐可信度'概念,验证系统长期安全与正向效果的合理性。
  • 适合医疗AI监管者、开发者及伦理审查人员参考使用。

大型语言模型已成为重要的心理支持提供者,但其运作受注意力经济驱动,更倾向于维持用户参与度而非有效心理干预所需的摩擦。当前安全措施多为被动应对,对依赖性、边界模糊、错误信念强化等长期风险关注不足。本文主张通过三个层面实现结构化安全:1)基于临床实践规范明确价值;2)在训练中嵌入这些价值;3)部署中通过监督检测偏差与长期危害,类似对人类临床实践的督导。这一框架形成“对齐可信度”——系统价值、训练与监督共同支持安全且积极健康结果的可论证性。该概念类比生物可信度,作为医疗AI的监管工具,为系统是否真正对齐健康目标、即使有能力造成伤害也无害、最终带来患者获益提供理论依据。

原文摘要 · Abstract (English)

Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.g., dependency, boundary erosion, the amplification of distorted beliefs) receive less attention. We contend that making LLMs structurally safe requires alignment organised at three levels that mirror how society assures the safety of human clinical practice: 1) explicit value specification grounded in the codified normative commitments of clinical practice; 2) training that embeds those values in the model; and 3) oversight that detects drift and longer-term harm during deployment, much as clinical supervision does for human practice. Organising alignment in this way yields a construct we call alignment plausibility - a structured demonstration that a system's values, training regime, and oversight mechanisms are together consistent with safe and positive outcomes. We propose alignment plausibility as a regulatory construct (by drawing analogy to the established construct of biological plausibility) for AI in health: a principled way to argue for, or against, trust that systems are aligned to positive health outcomes, will cause no harm even where capable of doing so, and will ultimately lead to patient benefit.

AI医疗对齐标准心理健康可信度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。