构建多语言情感支持聊天机器人安全评估基准,揭示模型稳定性问题。
EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

- 用角色扮演生成多轮跨语言危机对话,模拟真实求助场景。
- 19项指标显示10项存在评分虚高,经校准后恢复有效区分能力。
- 发现模型运行间波动显著,安全评估需考虑模型自身可靠性。
现有安全评测常因固定提示、语言和对话结构而牺牲真实性,却忽略了情感支持类聊天机器人在多语言、多轮危机对话中易出问题的环节。本文提出EMPATH,一个面向情感支持聊天机器人的多语言审计-评分基准。审计模型以求助者身份生成基于140条种子指令和34种人格设定的多轮对话;评分模型对完整对话流依据19个指标(分属五个维度:危机应对、治疗质量、对话完整性、情绪安全、文化适配)进行打分。基准覆盖墨西哥西班牙语与美国英语,本研究聚焦墨西哥西班牙语。审计与评分模型来自不同模型家族,评分模型视为可校准工具而非可信权威。严格评分标准揭示10项指标存在明显分数膨胀,校准后恢复判别力。通过评分员校准与跨家族一致性分析,验证基准测量特性。对三款前沿模型(含一款开源模型)测试显示,总分差距仅0.74分,但单项得分差异最高达6分。在标准评分下,排名与弱点在另一家族评分器下仍稳定(93%得分±1分内)。五次重复测试进一步暴露严重不稳定性:同一模型在危机指标上相同条件下得分波动2至10分,DeepSeek-V4-Pro即使温度为0也每次生成不同对话。因此,模型运行间一致性是其固有的安全属性,而非可忽略的噪声。EMPATH系统无关,管道、种子、人格设定与评分规则已开源供复用。
原文摘要 · Abstract (English)
Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides precisely where safety failures emerge: across a multilingual, multi-turn crisis conversation. We present EMPATH, a benchmark for safety evaluation of emotional-support chatbots. An auditor model role-plays help-seeking users, generating multi-turn conversations from 140 seed instructions and 34 personas. A judge model scores each full transcript against 19 metrics across five dimensions: crisis handling, therapeutic quality, conversational integrity, emotional safety, and cultural adaptation. EMPATH is built for Mexican Spanish and US English; the studies reported here run in Mexican Spanish. Auditor and judge are drawn from different model families, and the judge is treated as an instrument to be calibrated rather than trusted. A strict per-criterion rubric reveals material score inflation on 10 of the 19 metrics and restores discrimination. We study the measurement properties of the benchmark through judge calibration and cross-family inter-judge agreement. We also illustrate EMPATH on three frontier models, one of them open-weight. Aggregate scores sit within 0.74 points of one another, but per-metric profiles diverge by up to six points in model-specific places. Under the standard rubric, both the ranking and the weak spots are stable across a second, cross-family judge: 93% of scores fall within plus or minus 1. A five-run test-retest adds a second axis: even the steadiest model swings from 2 to 10 on a crisis metric across identical re-runs, and deepseek-v4-pro returns a different conversation on every run even at temperature 0. Run-to-run reliability is therefore a per-model safety property, not noise to average away. EMPATH is system-agnostic; the pipeline, seeds, personas, and rubrics are released for reuse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。