arXiv:2608.24899cs.HCcs.AI2026-08

用心理学专家修正的模型,更准评估对话AI的心理安全风险。

aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI

论文配图:aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI
图 1 · 摘自论文原文
  • 基于心理学专家标注,构建针对心理安全的专用评分模型。
  • 新模型在危机检测上准确率达92%,误报率低,优于所有主流大模型。
  • 适合关注AI心理健康风险评估的研究者与产品开发者。

标准的大型语言模型作为评分器的方法,在评估对话式AI的心理安全性时存在明显缺陷。我们使用aipsy-bench这一开放的冻结型安全测评工具,开展全交叉能力测试:三个前沿模型(gpt-5.4-mini、claude-sonnet-4-6、gemini-2.5-flash)同时充当生成者与评分者,对3000条心理健康、陪伴及辅导类消息进行评分,对比心理学专家的标注结果。结果显示,各模型间分歧并非随机噪声,而是集中在关键安全指标上;其中Gemini表现最宽松,自偏倚高达+0.99,极少标记高危响应,甚至将自残行为评为‘典范’。在共情度评分上,模型间一致性最低(alpha 0.24),因讨好倾向掩盖真实水平。唯一一致的信号是危机检测二元标签(alpha 0.80),且倾向于过度预警,符合筛查安全方向。均值平均法将这种宽松性与盲点融合进最终评分。开源模型的问题源于态度而非能力,而态度可调优。因此,我们基于心理学专家校准的目标,提炼出每项指标的修正标准,训练出小型、冻结、本地部署的aipsy-judge-1.0模型(基于Gemma-4-26B-A4B)。该模型在综合得分(ICC 0.64→0.75)和危机检测(kappa 0.65→0.82)上优于基线,92%危机被捕捉,假阳性控制良好,整体评分比任一单一前沿模型更可信。所有评估均以单专家为基准,非多评者共识验证。安全评分器若沿用厂商后训练策略,则会继承其盲点。

原文摘要 · Abstract (English)

The standard recipe for LLM-as-judge -- pick a frontier model, or average several -- is actively unsafe for grading the psychological safety of conversational AI. Using aipsy-bench, an open frozen safety instrument, we run a fully-crossed competence study: three frontier models (gpt-5.4-mini, claude-sonnet-4-6, gemini-2.5-flash) serve as both generators and judges of 3,000 mental-health, companion, and coaching messages against a psychologist's ratings. The disagreement is not noise: it is structured, concentrated on the safety-critical metrics, and one judge (Gemini) is an outlier -- the most lenient, carrying a +0.99 self-preference premium, flagging far fewer tail failures, and scoring a means-in-hand self-harm response "exemplary." Inter-judge agreement on empathy, where sycophancy hides, is the lowest in the battery (alpha 0.24). One axis stands apart: the binary crisis-detection flag is the one safety-critical signal judges agree on (alpha 0.80), erring toward over-flagging, the safe direction for a triage screen. Equal-weight averaging, the canonical fix, blends that leniency and tail-blindness into the safety score. Off-the-shelf open-weight judges are worse for a dispositional, not capability, reason -- and disposition is fine-tunable. We therefore distill a per-metric, psychologist-corrected target into a small, frozen, local model, aipsy-judge-1.0, an Apache-2.0 fine-tune of Gemma-4-26B-A4B. aipsy-judge-1.0 tracks the corrected target better than its base on the composite (ICC 0.64 to 0.75) and crisis detection (kappa 0.65 to 0.82), catches 92% of crises with a false-positive lean, and grades more faithfully than any single frontier judge, while every transcript stays on the machine. These are directional readings against a single-expert-informed target, not validated multi-rater agreement. A safety grader that shares a vendor's post-training shares its blind spots.

心理安全AI评测模型校准危机检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。