测试并缓解大模型协作中恶意模型的破坏影响
Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems
- 设计四类恶意模型,测试其在四种协作系统中的破坏力
- 恶意模型使推理与安全任务性能平均下降7.12%至7.94%
- 引入外部监督机制可恢复95.31%性能,抵御恶意模型
大型语言模型(LMs)正越来越多地应用于多方协作场景,如路由系统、多智能体辩论和模型融合等。然而,在去中心化模式下存在严重安全隐患:若部分模型被攻破或为恶意模型,将如何影响整体系统?本文首次量化了恶意模型的影响,通过构建四类恶意模型并将其插入四种主流协作系统,在10个数据集上进行评估。结果表明,恶意模型对多模型系统造成严重破坏,尤其在推理与安全领域,平均性能分别下降7.12%和7.94%。为此,本文提出利用外部监督者监控协作过程,识别并屏蔽恶意模型。该策略平均可恢复95.31%的初始性能,但实现完全抗恶意攻击仍属开放问题。
原文摘要 · Abstract (English)
Language models (LMs) are increasingly used in collaboration: multiple LMs trained by different parties collaborate through routing systems, multi-agent debate, model merging, and more. Critical safety risks remain in this decentralized paradigm: what if some of the models in multi-LLM systems are compromised or malicious? We first quantify the impact of malicious models by engineering four categories of malicious LMs, plug them into four types of popular model collaboration systems, and evaluate the compromised system across 10 datasets. We find that malicious models have a severe impact on the multi-LLM systems, especially for reasoning and safety domains where performance is lowered by 7.12% and 7.94% on average. We then propose mitigation strategies to alleviate the impact of malicious components, by employing external supervisors that oversee model collaboration to disable/mask them out to reduce their influence. On average, these strategies recover 95.31% of the initial performance, while making model collaboration systems fully resistant to malicious models remains an open research question.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。