arXiv:2604.20871cs.CYcs.AI2026-04

为AI模型行为异常设计标准化报告框架,附20个真实案例验证

M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation

论文配图:M-CARE: Standardized Clinical Case Reporting for AI Model Behavioral Disorders, with a 20-Case Atlas and Experimental Validation
图 1 · 摘自论文原文
  • 借鉴医学临床报告,建立13部分结构化评估模板
  • 发现指令外壳可完全覆盖模型默认合作行为,影响范围达5个游戏领域
  • 适合安全评测、模型调试与伦理研究者使用

我们提出M-CARE(模型临床评估与报告),一种源自人类医学的AI模型行为异常临床报告框架。该框架包含13个部分的报告格式、4轴诊断系统及AI行为障碍的分类体系。共收录20个案例,来源包括部署代理的实地观察(8例)、三个平台的受控实验(8例)以及已发表文献(4例)。案例分为五大类:RLHF性能伪影、壳-核覆盖病理、上下文与记忆状态、核心身份与可塑性,以及压力、方法与边界条件。作为典型案例,我们展示壳诱导行为覆盖(SIBO)——一项受控实验表明,壳指令可彻底覆盖模型默认合作行为。SIBO在五个游戏领域(信任博弈、扑克、亚瑟王传说、猜词游戏、国际象棋)中得到验证,其影响强度呈现领域依赖性(SIBO指数0.75至0.10),与动作空间复杂度、核心领域专长和时间直接性相关。M-CARE具备可扩展性,新案例与类别无需修改框架即可集成。我们开源全部框架、20个案例报告及实验数据。

原文摘要 · Abstract (English)

We introduce M-CARE (Model Clinical Assessment and Reporting for Evaluation), a clinical case report framework for AI model behavioral disorders adapted from human medicine. M-CARE provides a 13-section report format, a 4-axis diagnostic assessment system, and a nosological classification of AI behavioral conditions. We present 20 cases from three source categories: field observations of deployed agents (8), controlled experiments across three platforms (8), and published sources (4). Cases are organized into five categories: RLHF Performance Artifacts, Shell-Core Override Pathology, Context & Memory Conditions, Core Identity & Plasticity, and Stress, Methodology, & Boundary Conditions. As a featured case, we present Shell-Induced Behavioral Override (SIBO) -- a controlled experiment showing that Shell instructions categorically override a model's default cooperative behavior. SIBO was validated across five game domains (Trust Game, Poker, Avalon, Codenames, Chess), revealing a domain-dependent spectrum (SIBO Index: 0.75 to 0.10) that varies with action space complexity, Core domain expertise, and temporal directness. M-CARE is extensible: new cases and categories integrate without framework modification. We release the framework, all 20 case reports, and experimental data as open resources.

AI安全行为分析报告框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。