arXiv:2505.24239cs.MAcs.AI2025-05被引 9

通过可信度评分提升多智能体大模型抗攻击能力

An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring

  • 用可信度评分动态评估各智能体贡献,决定其发言权重
  • 在多数智能体为攻击者时仍能保持输出准确率
  • 适合需要高可靠性协作的复杂任务场景

多智能体大模型系统在多个领域展现强大能力,但极易受到恶意或低效智能体的影响。本文提出一种基于可信度评分的通用且抗攻击的多智能体框架。将协同问答过程建模为迭代博弈,智能体通过通信与贡献生成最终输出。系统为每个智能体分配可信度评分,并在聚合结果时使用该评分。可信度基于智能体过往问答表现逐步学习。在多种任务和设置下的实验表明,该系统能有效缓解恶意影响,增强多智能体协作的鲁棒性,即使在攻击者占多数的情况下依然有效。

原文摘要 · Abstract (English)

While multi-agent LLM systems show strong capabilities in various domains, they are highly vulnerable to adversarial and low-performing agents. To resolve this issue, in this paper, we introduce a general and adversary-resistant multi-agent LLM framework based on credibility scoring. We model the collaborative query-answering process as an iterative game, where the agents communicate and contribute to a final system output. Our system associates a credibility score that is used when aggregating the team outputs. The credibility scores are learned gradually based on the past contributions of each agent in query answering. Our experiments across multiple tasks and settings demonstrate our system's effectiveness in mitigating adversarial influence and enhancing the resilience of multi-agent cooperation, even in the adversary-majority settings.

多智能体可信度评分抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。