arXiv:2605.08827cs.AI2026-05被引 1

评估心理AI安全需保留对话时间线索,否则结论可能失效。

Mental Health AI Safety Claims Must Preserve Temporal Evidence

论文配图:Mental Health AI Safety Claims Must Preserve Temporal Evidence
图 1 · 摘自论文原文
  • 提出时间安全不可识别性,说明忽视对话顺序会误判安全
  • 在AnnoMI数据集上发现逐轮评分忽略的渐进式失败机制
  • 建议采用保留时间证据的评估标准,适用于临床级AI

当前心理AI的安全评估多基于孤立回应、最终结果或对话整体质量,但临床上的关键问题常源于交互的顺序与累积效应,如延迟升级、重复强化、依赖形成、修复失败及逐轮恶化。本文指出,这种评估尺度的错配不仅是覆盖不足,更是导致安全结论无效的根本原因。提出“时间安全不可识别性”概念,阐明依赖序列、时间、累积或恢复的安全属性无法通过丢弃这些特征的评估协议来验证。由此发展出通用原则SCOPE(安全声明需保留证据),并具体化为心理AI专用标准SCOPE-MH。通过在专家标注的心理访谈数据集AnnoMI上的原型验证,揭示了逐轮评分无法捕捉的故障机制。主张将SCOPE-MH作为现有评估体系的诊断补充,强调保留时间证据对安全关键型心理AI部署是必需而非可选。

原文摘要 · Abstract (English)

The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, or aggregate dialogue quality, while clinically consequential failures may arise from the order and accumulation of interactions themselves, including delayed escalation, repeated reinforcement, dependency formation, failed repair, and gradual deterioration across turns. This paper argues that this mismatch is not merely a limitation of evaluation coverage but a source of invalid safety conclusions. We introduce Temporal Safety Non-Identifiability, a formal account of why safety properties that depend on sequence, timing, accumulation, or recovery cannot be certified by protocols that discard those features. From this formalization, we develop SCOPE (Safety Claims Over Preserved Evidence) as a general principle for aligning safety claims with the evidence an evaluation actually retains, and instantiate it as SCOPE-MH, a mental-health instantiation of this reporting standard. We operationalize SCOPE-MH through a proof-of-concept on the AnnoMI dataset of expert-annotated motivational interviewing conversations, which reveals mechanisms of failure that per-turn behavior scoring does not represent. We propose SCOPE-MH as a diagnostic complement to existing evaluation infrastructure and argue that evaluation preserving temporal evidence is necessary, not optional, for safety-critical mental health AI deployment.

心理AI安全评估时间序列评测标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。