在预算有限时,如何用不可靠的模型自信度有效审计大量AI代理。
One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence

- 基于高斯耦合模型,分析信心排序审计在误差相关下的失效边界。
- 发现当信心偏差超过阈值δ*,按信心排序反而比随机审计更差。
- 实测五款开源模型信心几乎恒定,而一款闭源模型表现可信。
单个审核者需在每轮预算B远小于代理数量N的情况下,审计N个大语言模型代理。代理报告的信心可能被恶意扭曲且存在相关性错误。本文将此建模为两层高斯耦合的预算内噪声检测问题,确定了信心排序审计劣于随机审计的临界阈值δ*。两个直觉反转:δ*随预算缩减而上升;跨家族相关性不低,共性难度主导。对五款开放权重模型的测试显示其信心近乎恒定(操作上无意义),点估计或超出翻转点但置信区间跨越该点;而一款专有模型信心具信息量且低于阈值δ*。本文给出可操作的‘空洞监督’量化标准,并通过回放历史策略验证了审计顺序的有效性。
原文摘要 · Abstract (English)
A single human must audit $N$ LLM agents under a budget of $B \ll N$ audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the miscalibration threshold $δ^*$ past which confidence-ranked auditing is \emph{worse} than random. Two a-priori expectations reverse: $δ^*$ \emph{rises} as the budget shrinks, and cross-family correlation is not low---shared difficulty dominates lineage. Five open-weight LLMs show operationally useless (near-constant) confidence, point estimates at or beyond the flip though CIs straddle it; a proprietary model is informative and lands below it. We give a quantitative criterion for \emph{vacuous} oversight, and replaying policies on recorded traces confirms the ordering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。