arXiv:2606.14200cs.AIcs.LG2026-06

针对智能体群组的技能依赖信任机制,揭示其优势与安全漏洞。

When Should Agent Trust Be Conditional? Characterizing and Attacking Skill-Conditional Reputation in Agent Swarms

论文配图:When Should Agent Trust Be Conditional? Characterizing and Attacking Skill-Conditional Reputation in Agent Swarms
图 1 · 摘自论文原文
  • 用技能条件信任取代单一全局评分,提升任务匹配精度
  • 在高异质性、低证据密度场景下,条件信任可带来真实性能增益
  • 攻击者可利用跨技能信息伪装,暴露信任机制的安全风险

开放平台日益在异构大模型智能体间分配任务,这些智能体在基础模型、架构和工具栈上差异显著,其能力随技能变化剧烈。传统全局信任评分无法体现专业化价值。本文研究技能条件信任 R(i|k)——对执行某技能任务的智能体 i 的信任度,而非单一评分,并提出三个可验证问题:条件信任何时值得采用、应借多少跨技能证据、该借用是否安全。通过受控相图分析发现,条件信任仅在高智能体异质性、单技能证据稀疏且技能相关时有效;耦合强度 beta 虽提升数据效率,但亦成为信息清洗通道。在包含14个真实异构AppWorld智能体的公开基准测试中,实际群体位于有利区间,实现小幅但真实的性能提升,且各技能最优智能体确实变动。进一步实验表明,攻击者仅需少量目标技能外的低成本证据,即可劫持条件路由系统,使路由损失从0飙升至0.94;而我们提出的零成本条件信息价值测试(CIVT)仍判定为绿色,原信任判据却由诚实的+0.19被篡改为-0.06。零证据门限虽可限制攻击,但无法根除,我们明确刻画了在显式预算下的残留代价。本文不主张完全抗伪造,而是量化信任与风险之间的权衡。

原文摘要 · Abstract (English)

Open platforms increasingly route tasks among heterogeneous LLM agents--differing in base model, scaffold, and tool stack--whose competence varies sharply by skill: an agent excellent at one skill may be useless at another. The standard reputation approach summarizes each agent by a single global trust score, but that scalar is the wrong object here, because routing every task to the globally most-trusted agent leaves the value of specialization unclaimed. We study skill-conditional trust R(i | k)--the trust to place in agent i for a task requiring skill k, rather than one score per agent--and pose three falsifiable questions: when is conditioning worth it, how much cross-skill evidence should be borrowed, and whether that borrowing is safe. A controlled phase-diagram analysis answers the first two: conditional trust wins only in a specific regime--high agent heterogeneity, sparse per-skill evidence, and correlated skills--and the coupling strength beta that buys this data efficiency is dual-use, because the same cross-skill borrowing is also a laundering channel. On a public benchmark of 14 genuinely heterogeneous AppWorld agents, real pools land inside the beneficial regime--a small but genuine gain, with the per-skill best agent genuinely changing across skills. We then show that an attacker with cheap evidence in one skill and none in a target skill hijacks the conditional router, driving routing regret from 0 to 0.94 on a pool our zero-cost Conditional Information Value Test (CIVT) rates GREEN--while the ungated trust verdict it contaminates reads -0.06 instead of the honest +0.19. A zero-evidence gate bounds the attack but does not eliminate it; we characterize the residual cost under an explicit budget. We do not claim Sybil-resistance--we quantify the trade-off.

智能体协作信任机制安全攻防多技能调度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。