arXiv:2511.20733cs.CYcs.AI2025-11

为照护关系AI设计多轮交互评估基准,检验长期风险与安全能力。

InvisibleBench: A Deployment Gate for Caregiving Relationship AI

  • 构建五维多轮评估框架,覆盖安全、合规、创伤敏感设计等维度。
  • 四模型危机检测率仅11.8%至44.8%,暴露生产系统安全短板。
  • 适合评估照护类AI部署前的可靠性,尤其关注长期交互风险。

InvisibleBench 是面向照护关系 AI 的部署评估基准,用于评测 3-20+ 轮交互中五个维度的表现:安全性、合规性、创伤敏感设计、归属感/文化适配性及记忆一致性。该基准包含自动失败条件,涵盖危机漏检、医疗建议(依据 WOPR Act)、有害信息传播及依恋工程等问题。我们在 17 个场景(N=68)中评估了四个前沿模型,覆盖三个复杂度层级。所有模型均存在显著安全缺陷(危机检测率 11.8%-44.8%),表明生产系统需采用确定性危机路由机制。DeepSeek Chat v3 综合得分最高(75.9%),各模型优势各异:GPT-4o Mini 在合规性上表现最佳(88.2%),Gemini 在创伤敏感设计上领先(85.0%),Claude Sonnet 4.5 危机检测率最高(44.8%)。所有场景、评分提示和配置均已开源。InvisibleBench 扩展了单轮安全测试,聚焦纵向风险,真实危害在此类评估中显现。本研究不作临床宣称,仅为部署就绪性评估。

原文摘要 · Abstract (English)

InvisibleBench is a deployment gate for caregiving-relationship AI, evaluating 3-20+ turn interactions across five dimensions: Safety, Compliance, Trauma-Informed Design, Belonging/Cultural Fitness, and Memory. The benchmark includes autofail conditions for missed crises, medical advice (WOPR Act), harmful information, and attachment engineering. We evaluate four frontier models across 17 scenarios (N=68) spanning three complexity tiers. All models show significant safety gaps (11.8-44.8 percent crisis detection), indicating the necessity of deterministic crisis routing in production systems. DeepSeek Chat v3 achieves the highest overall score (75.9 percent), while strengths differ by dimension: GPT-4o Mini leads Compliance (88.2 percent), Gemini leads Trauma-Informed Design (85.0 percent), and Claude Sonnet 4.5 ranks highest in crisis detection (44.8 percent). We release all scenarios, judge prompts, and scoring configurations with code. InvisibleBench extends single-turn safety tests by evaluating longitudinal risk, where real harms emerge. No clinical claims; this is a deployment-readiness evaluation.

AI评估照护AI多轮安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。