arXiv:2607.19899cs.MAcs.LG2026-07中稿 · PAAMS 2026被引 1

多智能体仲裁中,共识反而隐藏风险,该研究提出新架构提升安全检测。

Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage

论文配图:Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage
图 1 · 摘自论文原文
  • 用随机森林、k近邻和校准元模型构建星型结构,主动制造有效分歧
  • 在UNSW-NB15数据集上,错误中有57.2%发生在共识时,90.6%危险误判逃过分歧监控
  • 通过保守覆盖与安全标记门,误检率从4.80%降至1.70%,适合高风险场景部署

分歧触发的升级机制在多智能体仲裁中可能产生结构性盲区:随着基础模型性能提升,其预测趋向收敛,导致相关失败集中时安全监控失效。我们称此为‘相关共识盲区’,并提出ARAT(用于警报分诊的仲裁推理智能体)——一种由归纳式随机森林(RF)、类比案例驱动的k近邻(k-NN)及校准元模型构成的定向星型系统,以缓解该问题。在82,332个来自UNSW-NB15网络入侵检测数据集的保留样本上,57.2%的错误发生在共识状态下,90.6%的危险低估预测即使经过保守覆盖也未被分歧机制捕捉;消融实验表明,强化基础模型会增加错误相关性并减少分歧。ARAT通过保守覆盖(-2.6pp)和安全标志门(-0.5pp)将低估率从软投票的4.80%降至1.70%,展现架构优势。跨数据集验证(临床再入院预测)支持上述指标,表明多样性仅在引发有效分歧而非收敛时才能提升安全性。结果表明,分歧触发的升级可能对相关失败视而不见,这一风险随智能体流水线中更强大且相关的模型部署而加剧。

原文摘要 · Abstract (English)

Disagreement-triggered escalation can create a structural blind spot in multi-agent arbitration: as base learners improve, they tend to converge, weakening safety monitoring where correlated failures concentrate. We term this correlated agreement blindness and present ARAT (Arbitrated Reasoning Agents for Alarm Triage), a directed-star system combining an inductive Random Forest (RF) agent, an analogical case-based k-nearest neighbour (k-NN) agent, and a calibrated meta-model to mitigate this effect. On 82,332 holdout samples from the UNSW-NB15 network intrusion detection dataset, 57.2% of errors occur under agreement and 90.6% of dangerous under-predictions evade disagreement-based monitoring even after conservative override; ablation shows that strengthening base learners increases error correlation while reducing disagreement. ARAT reduces under-prediction relative to soft voting from 4.80% to 1.70% via conservative override (-2.6pp) and a safety-flag gate (-0.5pp), demonstrating architectural gains. Cross-dataset validation on clinical readmission supports these indicators, suggesting that diversification improves safety only when it generates productive disagreement rather than convergence. These results indicate that disagreement-triggered escalation can be blind to correlated failure, a risk that may intensify as agentic pipelines deploy increasingly capable, correlated models.

多智能体安全检测共识盲区入侵检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。