arXiv:2512.02282cs.AIcs.HC2025-12被引 5

用多智能体评估大模型回应的心理社会风险,提升敏感场景安全性。

DialogGuard: Multi-Agent Psychosocial Safety Evaluation of Sensitive LLM Responses

  • 构建多智能体框架,通过四种评分机制检测五类高危行为。
  • 双智能体修正和多数投票法在准确率与人类判断一致性上最优。
  • 开源工具支持可解释性分析,适合心理健康类应用开发者使用。

大型语言模型(LLMs)如今广泛用于网络心理卫生、危机干预等情感敏感服务,但其在这些场景中的心理社会安全性仍不明确且评估薄弱。我们提出 DialogGuard,一种多智能体框架,用于评估 LLM 生成回应在隐私侵犯、歧视行为、心理操控、心理伤害和侮辱性言论五个高严重性维度上的风险。DialogGuard 可通过四种基于 LLM 作为评判者的管道应用到多种生成模型:单智能体评分、双智能体修正、多智能体辩论和随机多数投票,均基于一个供人类标注者和 LLM 判官共用的三级评分标准。利用带人工安全标注的 PKU-SafeRLHF 数据集,我们发现多智能体机制比非 LLM 基线和单智能体判断更准确;双智能体修正与多数投票在准确性、与人类评分的一致性及鲁棒性之间取得最佳平衡,而辩论虽召回率更高,但对边界案例过警报。我们开源了 DialogGuard 软件,配备网页界面,提供各维度风险评分及可解释的自然语言理由。12 名从业者的形成性研究显示,该工具可有效支持提示设计、审计与面向脆弱用户的在线应用监管。

原文摘要 · Abstract (English)

Large language models (LLMs) now mediate many web-based mental-health, crisis, and other emotionally sensitive services, yet their psychosocial safety in these settings remains poorly understood and weakly evaluated. We present DialogGuard, a multi-agent framework for assessing psychosocial risks in LLM-generated responses along five high-severity dimensions: privacy violations, discriminatory behaviour, mental manipulation, psychological harm, and insulting behaviour. DialogGuard can be applied to diverse generative models through four LLM-as-a-judge pipelines, including single-agent scoring, dual-agent correction, multi-agent debate, and stochastic majority voting, grounded in a shared three-level rubric usable by both human annotators and LLM judges. Using PKU-SafeRLHF with human safety annotations, we show that multi-agent mechanisms detect psychosocial risks more accurately than non-LLM baselines and single-agent judging; dual-agent correction and majority voting provide the best trade-off between accuracy, alignment with human ratings, and robustness, while debate attains higher recall but over-flags borderline cases. We release Dialog-Guard as open-source software with a web interface that provides per-dimension risk scores and explainable natural-language rationales. A formative study with 12 practitioners illustrates how it supports prompt design, auditing, and supervision of web-facing applications for vulnerable users.

心理安全多智能体LLM评估敏感对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。