arXiv:2602.00950cs.AI2026-02被引 3

专为心理健康对话设计的安全防护模型,能精准识别危机与非危机内容。

MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support

  • 基于临床心理学构建风险分类体系,区分真实危机与正常倾诉。
  • 在高召回率下降低误报率,对抗攻击时有害互动率更低。
  • 适合心理支持类大模型开发者使用,提升对话安全性。

大型语言模型在心理健康支持中应用日益广泛,但其对话连贯性不足以保证临床恰当性。现有通用安全机制难以区分治疗性倾诉与真实临床危机,导致安全失效。为此,我们与博士级心理学家合作,构建了临床基础的风险分类体系,可识别自伤、伤人等可行动危害,同时保留安全的非危机治疗内容空间。我们发布了MindGuard-testset,一个由临床专家逐轮标注的真实多轮对话数据集。通过受控双智能体生成的合成对话,训练了4B和8B参数的轻量级安全分类器(MindGuard)。该分类器在高召回率下显著降低误报率;与临床语言模型结合后,在对抗性多轮交互中,攻击成功率与有害互动率均低于通用防护方案。所有模型及人工评估数据均已开源。

原文摘要 · Abstract (English)

Large language models are increasingly used for mental health support, yet their conversational coherence alone does not ensure clinical appropriateness. Existing general-purpose safeguards often fail to distinguish between therapeutic disclosures and genuine clinical crises, leading to safety failures. To address this gap, we introduce a clinically grounded risk taxonomy, developed in collaboration with PhD-level psychologists, that identifies actionable harm (e.g., self-harm and harm to others) while preserving space for safe, non-crisis therapeutic content. We release MindGuard-testset, a dataset of real-world multi-turn conversations annotated at the turn level by clinical experts. Using synthetic dialogues generated via a controlled two-agent setup, we train MindGuard, a family of lightweight safety classifiers (with 4B and 8B parameters). Our classifiers reduce false positives at high-recall operating points and, when paired with clinician language models, help achieve lower attack success and harmful engagement rates in adversarial multi-turn interactions compared to general-purpose safeguards. We release all models and human evaluation data.

心理健康安全防护多轮对话分类器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。