提出可审计的评估方法,精准区分路由效果与答案差异。
COVER: Identifiable Evaluation of Coalition Routing
- 固定信息边界和团队集合,通过对比不同策略组验证路由影响。
- 在多个数据集上显示路由改进显著,但原始答案提升微弱且不一致。
- 适合关注多智能体系统评估公平性与可解释性的研究者使用。
当多智能体系统更换团队时,其传递的消息和最终答案也会改变,因此端到端准确率差距无法独立识别路由效应。本文提出一种评估契约(COVER),在结果生成前固定公共信息边界、下游任务栈G及有限合法团队族。完全覆盖可识别在该栈条件下的精确有限基准奥拉克后悔值。对于任意有限冻结策略集合,执行其不同团队的并集是进行任意两两策略对比的最小无假设支持,但不适用于绝对奥拉克后悔。两个受控实验测试该工具:在MuSiQue-12上,预设特权正向控制使后悔值从0.532降至0.402;后期公开接口控制达到0.424(对比0.554),但为回溯性。在HotpotQA-4上,预设公开直接评分器将后悔值从0.313降至0.110。固定栈下Llama执行中,验证路由后悔改善0.190,而原始答案提升仅0.010且置信区间跨越零。五家族ToolSandbox变体验证对14个未触碰任务变体(224/224有效行)进行了全面评估:声明家族奥拉克达到0.768安全证据完成率,前瞻性冻结路由器为0.637(后悔0.131),未达预设0.10标准。后续回溯比较器达0.655,与所有工作者平均4.57对5.00工人相当。因此,COVER暴露了选择余地,而非虚构路由优势。交叉栈诊断显示绝对得分依赖于G,但未发现可检测的路由器与终结器交互作用。COVER是一种可审计的测量方法,非堆栈不变或通用智能体路由优越性的主张。
原文摘要 · Abstract (English)
When a multi-agent system changes its team, it also changes the messages and final answer it produces, so an end-to-end accuracy gap does not by itself identify a routing effect. We introduce method, an evaluation contract that fixes a public information boundary, downstream stack G, and finite legal team family before outcomes are generated. Complete coverage identifies exact finite-benchmark oracle regret conditional on that stack. For any finite collection of frozen policies, executing the union of their distinct selected teams is the minimal assumption-free support for every pairwise policy contrast, though not for absolute oracle regret. Two controlled tables with source-ID-disjoint splits test the instrument. On MuSiQue-12, a pre-specified privileged positive control improves regret from 0.532 to 0.402; a later public-interface control reaches 0.424 versus 0.554 but is retrospective. On HotpotQA-4, a pre-specified public direct scorer improves regret from 0.313 to 0.110. In fixed-stack Llama execution, verified route regret improves by 0.190, while the raw-answer gain is 0.010 with an interval crossing zero. A five-family ToolSandbox variant-shift validation exhaustively evaluates 16 declared teams on 14 untouched task variants (224/224 valid rows): the declared-family oracle reaches 0.768 safe-evidence completion, while the prospectively frozen router gets 0.637 (regret 0.131), failing the predeclared 0.10 criterion. A later retrospective comparator reaches 0.655, matching all-workers with 4.57 versus 5.00 workers on average. Thus COVER exposes selection headroom without manufacturing a routing win. A crossed-stack diagnostic shows absolute scores depend on G but finds no detectable router-by-finalizer interaction. COVER is an auditable measurement methodology, not a claim of stack-invariant or universal agent-routing superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。