arXiv:2601.12754cs.HCcs.AI2026-01被引 7

用双代理框架实时审核并改进AI心理支持回复质量。

PAIR-SAFE: A Paired-Agent Approach for Runtime Auditing and Refining AI-Mediated Mental Health Support

  • 构建响应者与裁判者双代理,裁判基于临床标准评估回复
  • 实测显示合作性、关系质量等关键维度显著提升
  • 适合关注AI心理服务安全性的研究者与开发者

大型语言模型在心理健康支持中日益普及,但可能生成过于指令化、不一致或临床错位的回应,尤其在敏感或高风险情境下。现有缓解策略多依赖训练或提示中的隐式对齐,缺乏透明度和运行时问责。我们提出PAIR-SAFE,一种结合响应者与基于临床验证的动机访谈完整性(MITI-4)框架的裁判者的配对代理框架,用于实时审计与优化生成内容。裁判者对每条回复进行结构化判定(允许或修正),指导运行时优化。通过基于人类标注的动机访谈数据构建的支持者模拟器,我们模拟咨询交互。结果显示,经裁判监督的互动在伙伴关系、协作寻求及整体关系质量等关键MITI维度上均有显著提升。定量结果获专家定性评估支持,凸显运行时监督的细微价值。研究证明,该双代理方法能为人工智能辅助对话式心理支持提供临床基础的审计与优化能力。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used for mental health support, yet they can produce responses that are overly directive, inconsistent, or clinically misaligned, particularly in sensitive or high-risk contexts. Existing approaches to mitigating these risks largely rely on implicit alignment through training or prompting, offering limited transparency and runtime accountability. We introduce PAIR-SAFE, a paired-agent framework for auditing and refining AI-generated mental health support that integrates a Responder agent with a supervisory Judge agent grounded in the clinically validated Motivational Interviewing Treatment Integrity (MITI-4) framework. The Judgeaudits each response and provides structuredALLOW or REVISE decisions that guide runtime response refinement. We simulate counseling interactions using a support-seeker simulator derived from human-annotated motivational interviewing data. We find that Judge-supervised interactions show significant improvements in key MITI dimensions, including Partnership, Seek Collaboration, and overall Relational quality. Our quantitative findings are supported by qualitative expert evaluation, which further highlights the nuances of runtime supervision. Together, our results reveal that such pairedagent approach can provide clinically grounded auditing and refinement for AI-assisted conversational mental health support.

心理支持双代理MITI运行时审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。