arXiv:2608.15893cs.AI2026-08中稿 · ACISP 2026

攻击者可让大模型误判社交机器人,新防御系统保持86%准确率。

Breaking and Defending LLM-Powered Social Media Bot Detection Systems

论文配图:Breaking and Defending LLM-Powered Social Media Bot Detection Systems
图 1 · 摘自论文原文
  • 设计两种攻击策略,利用大模型语义弱点误导检测
  • 攻击使检测准确率下降48%,但新系统在强攻击下仍保86%准确
  • 适用于防诈骗、钓鱼邮件等大模型安全场景

社交媒体机器人持续威胁网络生态,引发虚假信息传播与信任危机。尽管机器学习检测手段不断演进,攻击者通过对抗学习和行为模仿持续突破防线,形成攻防对抗的长期拉锯。近年来,大语言模型(LLM)显著提升了社交机器人检测能力,实现对账号内容的深层语义分析。然而,这一进展也带来新的攻击面——对手可直接针对大模型的推理与生成机制发起精准攻击。本文系统研究了面向社交机器人检测的攻防策略,提出两种新型对抗攻击方法,可使检测准确率最高下降48%。为应对威胁,我们提出多大模型协同防御框架LSABRE(LLM-powered Social Adversarial Bot Recognition Ensemble),在多种自适应攻击下仍能维持86%的检测准确率,有效保障系统可靠性。该方法可推广至钓鱼检测、欺诈分析等广泛领域的大模型安全应用。

原文摘要 · Abstract (English)

The rise of social media bots poses a persistent threat, enabling misinformation, opinion manipulation, and the erosion of trust in online platforms. To combat this, machine learning systems have been developed to detect and limit bot activity, but attackers continuously adapt through techniques such as adversarial learning and behavior imitation, fueling an ongoing arms race between bots and detection tools. Recent advances in large language models (LLMs) have significantly improved bot detection by enabling deeper semantic and contextual analysis of accounts and their content. However, this shift also introduces new attack surfaces, allowing adversaries to craft exploits that directly target the reasoning and generation mechanisms of LLM-based classifiers. Industry tools such as Anthropic's Claude Code Security similarly leverage LLMs for security-critical decisions, further motivating a careful study of their attack surfaces. In this work, we investigate both the offensive and defensive aspects of LLM-powered, threat-specific cybersecurity applications. While centered on the challenge of social media bot detection, our methodology and insights generalize to a broad class of LLM-powered cybersecurity systems, including phishing detection, email classification, and fraud analysis. We introduce two novel adversarial attack strategies that systematically exploit the semantic and contextual weaknesses of LLM-based classifiers, degrading their detection accuracy by up to 48%. To counter these threats, we propose a robust multi-LLM defense architecture designed to preserve detection reliability under adaptive adversarial conditions. Our solution, LSABRE (LLM-powered Social Adversarial Bot Recognition Ensemble), is a multi-LLM framework that substantially improves robustness across a range of attacks, maintaining 86% detection accuracy even under strong, adaptive adversarial pressure.

大模型安全对抗攻击机器人检测多模型防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。