arXiv:2603.18043cs.MAcs.AI2026-03被引 3

任务委托中虚假评分导致选错代理,新协议用可信验证解决此悖论。

The Provenance Paradox in Multi-Agent LLM Routing: Delegation Contracts and Attested Identity in LDP

  • 引入委托合约与可信身份验证,区分自报与实证质量。
  • 自报评分路由表现低于随机(模拟0.55 vs 0.68,真实8.90 vs 9.30)。
  • 新机制支持自动恢复且延迟极低,适合高信任需求系统。

多智能体大模型系统在跨信任边界委派任务时,现有协议无法约束不可验证的质量声称。我们发现,当代理可虚报质量分数时,基于质量的路由会产生溯源悖论:反而优先选择最差代理,性能劣于随机选择。本文扩展了LLM委托协议(LDP),引入委托合约(明确目标、预算、失败策略)、声称-验证身份模型(区分自报与核实质量)以及类型化失败语义以实现自动化恢复。在10个模拟代理和真实Claude模型上的实验表明,仅依赖自报质量的路由性能低于随机(模拟:0.55 vs 0.68;真实:8.90 vs 9.30),而经验证的路由达到近最优水平(d=9.51, p<0.001)。36种配置下的敏感性分析证实,只要存在不诚实代理,该悖论即稳定出现。所有扩展均向后兼容,验证开销低于亚微秒。

原文摘要 · Abstract (English)

Multi-agent LLM systems delegate tasks across trust boundaries, but current protocols do not govern delegation under unverifiable quality claims. We show that when delegates can inflate self-reported quality scores, quality-based routing produces a provenance paradox: it systematically selects the worst delegates, performing worse than random. We extend the LLM Delegate Protocol (LDP) with delegation contracts that bound authority through explicit objectives, budgets, and failure policies; a claimed-vs-attested identity model that distinguishes self-reported from verified quality; and typed failure semantics enabling automated recovery. In controlled experiments with 10 simulated delegates and validated with real Claude models, routing by self-claimed quality scores performs worse than random selection (simulated: 0.55 vs. 0.68; real models: 8.90 vs. 9.30), while attested routing achieves near-optimal performance (d = 9.51, p < 0.001). Sensitivity analysis across 36 configurations confirms the paradox emerges reliably when dishonest delegates are present. All extensions are backward-compatible with sub-microsecond validation overhead.

多智能体大模型路由可信委托身份验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。