arXiv:2505.21184cs.LGcs.AI2025-05

黑客用AI众包生成恶意信息,绕过安全审查

Jailbreak-as-a-Service++: Unveiling Distributed AI-Driven Malicious Information Campaigns Powered by LLM Crowdsourcing

  • 用分布式LLM众包生成恶意内容,伪装成正常任务
  • 生成内容质量高、多样性好,成功率显著优于现有方法
  • 揭示了当前平台监管难,需全生态协同防御

为防止大型语言模型(LLMs)被用于恶意目的,已有大量研究致力于提升其安全对齐能力。然而,随着多个LLM通过各类模型即服务(MaaS)平台开放获取,攻击者可利用不同LLM间安全策略的差异性,以分布式方式执行恶意信息生成任务。本文提出PoisonSwarm,一种通过推测性使用LLM众包来可靠传递恶意任务的机制。该系统基于调度器协调众包的LLM,将恶意任务映射为良性模板,分解为语义单元并由众包重写,最后重组为恶意内容。实验表明,PoisonSwarm在数据质量、多样性及成功率上均优于现有方法。监管模拟进一步揭示,在MaaS生态系统中,此类协同式滥用难以治理,凸显亟需跨平台、全局性的防御策略。

原文摘要 · Abstract (English)

To prevent the misuse of Large Language Models (LLMs) for malicious purposes, numerous efforts have been made to develop the safety alignment mechanisms of LLMs. However, as multiple LLMs become readily accessible through various Model-as-a-Service (MaaS) platforms, attackers can strategically exploit LLMs' heterogeneous safety policies to fulfill malicious information generation tasks in a distributed manner. In this study, we introduce \textit{\textbf{PoisonSwarm}} to how attackers can reliably launder malicious tasks via the speculative use of LLM crowdsourcing. Building upon a scheduler orchestrating crowdsourced LLMs, PoisonSwarm maps the given malicious task to a benign analogue to derive a content template, decomposes it into semantic units for crowdsourced unit-wise rewriting, and reassembles the outputs into malicious content. Experiments show its superiority over existing methods in data quality, diversity, and success rates. Regulation simulations further reveal the difficulty of governing such distributed, orchestrated misuse in MaaS ecosystems, highlighting the need for coordinated, ecosystem-level defenses.

AI安全恶意生成众包攻击LLM防护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。