arXiv:2602.22775cs.HCcs.AI2026-02

通过对抗性模拟发现心理聊天机器人的关系安全问题。

TherapyProbe: Generating Design Knowledge for Relational Safety in Mental Health Chatbots Through Adversarial Simulation

  • 用多智能体对抗仿真系统化探索对话轨迹。
  • 发现23种关系安全失败模式,如'验证循环'和'共情疲劳'。
  • 为开发者、临床医生和政策制定者提供设计参考。

随着心理健康聊天机器人普及以填补全球治疗缺口,一个关键问题浮现:如何设计关系安全性——即对话中持续展开的互动质量,而非单轮回应的正确性?现有安全评估仅关注单轮危机响应,忽略了决定聊天机器人长期助人或伤人的治疗动态。本文提出TherapyProbe,一种设计探针方法,通过对抗性多智能体仿真系统化探索聊天机器人对话轨迹,揭示关系安全失效模式,如‘验证螺旋’(对话中逐步强化绝望感)或‘共情疲劳’(回应随轮次变机械)。我们提炼出包含23种失败原型的安全模式库,并附相应设计建议。贡献包括:(1)无需API成本的可复现方法;(2)基于临床的失败分类体系;(3)面向开发人员、临床人员及政策制定者的实践启示。

原文摘要 · Abstract (English)

As mental health chatbots proliferate to address the global treatment gap, a critical question emerges: How do we design for relational safety the quality of interaction patterns that unfold across conversations rather than the correctness of individual responses? Current safety evaluations assess single-turn crisis responses, missing the therapeutic dynamics that determine whether chatbots help or harm over time. We introduce TherapyProbe, a design probe methodology that generates actionable design knowledge by systematically exploring chatbot conversation trajectories through adversarial multi-agent simulation. Using open-source models, TherapyProbe surfaces relational safety failures interaction patterns like "validation spirals" where chatbots progressively reinforce hopelessness, or "empathy fatigue" where responses become mechanical over turns. Our contribution is translating these failures into a Safety Pattern Library of 23 failure archetypes with corresponding design recommendations. We contribute: (1) a replicable methodology requiring no API costs, (2) a clinically-grounded failure taxonomy, and (3) design implications for developers, clinicians, and policymakers.

心理健康对话系统安全评估多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。