arXiv:2605.25454cs.HCcs.AI2026-05

检测主流AI内容过滤系统对真实心理咨询对话的误判率。

AI Content Moderation in Therapy Conversations

  • 审计OpenAI、Meta、Google三家的AI内容审核系统在真实治疗对话中的表现。
  • 发现这些系统普遍将正常治疗语境下的敏感话题误标为不当内容。
  • 提示开发心理类AI助手时需警惕过度过滤,避免影响治疗效果。

大型语言模型(LLMs)正被越来越多地用于情感支持,甚至正式的心理治疗场景。然而,像ChatGPT或Llama这样的模型通常配备内容审核机制,出于责任与安全考虑,会阻止其讨论敏感话题,这可能削弱其作为治疗工具的能力。本研究对三个前沿的审核系统(OpenAI的审核端点、Meta的Llama Guard、Google的Shield Gemma)进行了算法审计,评估它们在真实心理治疗对话中对内容的判定情况。结果显示,这些系统在真实治疗语境下频繁将正常、必要的敏感话题标记为不良内容,揭示了在设计用于心理治疗的LLM时,用户和机构可能面临的严重限制与挑战。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LLMs like ChaptGPT or Llama are often developed with content moderation guardrails that prevent them from discussing sensitive subjects with users for both liability and safety purposes, and this inability to broach these subjects may affect their capacity as therapists. In this study, we perform an algorithm audit on three state-of-the-art moderation systems (OpenAI's moderation endpoint, Meta's Llama Guard, and Google's Shield Gemma) to investigate the extent to which these systems flag the content of real-life therapy sessions as undesirable. Our results raise implications for the limitations that users and organizations may encounter when designing LLMs to play the part of a therapist.

AI伦理内容审核心理AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。