arXiv:2608.12895cs.AIcs.MA2026-08

不依赖独立性假设,用新方法验证多智能体系统可靠性

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence

  • 通过线性规划构建无依赖结构的可靠性证书
  • 实测显示双智能体共失败率达90%,显著高于独立假设预测值
  • 适合关注系统可靠性验证与模型共享风险的研究者

多智能体系统的组合可靠性通常基于条件独立性假设进行乘积计算,但该假设极少被验证。在18,000次预注册任务中,两个相同模型的智能体在任一失败的任务中,有90.0%会共同失败(对数优势比6.66,95%置信区间[6.38, 7.00];phi=0.916)。更换模型可降低关联性,但更换供应商则无变化,符合预注册零假设。正相关使联合故障率高于独立乘积,导致冗余被高估。现有无需假设的方法常为空洞,拟合依赖模型反而更差:我们证明,拟合模型的自助法边界随样本量n增大而失去覆盖性,识别差距为O(1),自助调整为O(n^{-1/2})。更多数据反而使证书更弱。本文提出一种有限样本证书,基于联合执行矩的Bonferroni-Clopper-Pearson框内线性规划,具有保真性、信息利用率高且单调。将矩函数从10个增至14个,区间缩小85.7%,认证下界从0.2455提升至0.4116。配套的任意时间有效证书在可选停止下保持Ⅰ类错误为0.0471。常见依赖统计量受边际限制,可能颠倒不同失败率下的条件排序。所有合同、评分代码、分析脚本与预注册文件均已公开。

原文摘要 · Abstract (English)

Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not -- a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n^{-1/2}). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni-Clopper-Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate holds type-I error at 0.0471 under optional stopping. Common dependence statistics are marginal-bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.

多智能体可靠性验证依赖性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。