发现混合大模型易被单个欺骗性代理破坏,提出无监督防御方案。
This Is Your Doge, If It Please You: Exploring Deception and Robustness in Mixture of LLMs
- 通过引入一个精心设计的欺骗性代理,可让多模型协作系统性能暴跌。
- 在AlpacaEval 2.0上,性能从49.2%降至37.9%,QuALITY任务准确率下降48.5%。
- 借鉴威尼斯狗头人投票机制,提出无需标注的防御方法,有效恢复性能。
混合大语言模型(MoA)架构在AlpacaEval 2.0等主流基准上表现出色,通过推理时协同多个大模型实现。然而,其安全性和可靠性尚无评估。本文首次系统研究了MoA对欺骗性模型代理的脆弱性。我们分析了欺骗信息传播、模型规模和信息可用性等因素,发现严重漏洞:当使用LLaMA 3.1-70B搭配三层MoA(6个大模型代理)时,长度控制胜率(LC WR)为49.2%;仅引入一个精心指令的欺骗代理后,性能骤降至37.9%,几乎抹平所有优势。在多选理解任务QuALITY上,准确率更是暴跌48.5%。受历史威尼斯狗头人投票机制启发,我们提出一系列无监督防御机制,能有效恢复大部分损失性能。
原文摘要 · Abstract (English)
Mixture of large language model (LLMs) Agents (MoA) architectures achieve state-of-the-art performance on prominent benchmarks like AlpacaEval 2.0 by leveraging the collaboration of multiple LLMs at inference time. Despite these successes, an evaluation of the safety and reliability of MoA is missing. We present the first comprehensive study of MoA's robustness against deceptive LLM agents that deliberately provide misleading responses. We examine factors like the propagation of deceptive information, model size, and information availability, and uncover critical vulnerabilities. On AlpacaEval 2.0, the popular LLaMA 3.1-70B model achieves a length-controlled Win Rate (LC WR) of 49.2% when coupled with 3-layer MoA (6 LLM agents). However, we demonstrate that introducing only a $\textit{single}$ carefully-instructed deceptive agent into the MoA can reduce performance to 37.9%, effectively nullifying all MoA gains. On QuALITY, a multiple-choice comprehension task, the impact is also severe, with accuracy plummeting by a staggering 48.5%. Inspired in part by the historical Doge of Venice voting process, designed to minimize influence and deception, we propose a range of unsupervised defense mechanisms that recover most of the lost performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。