arXiv:2503.08704cs.CRcs.AI2025-03被引 11

揭示大模型路由在全生命周期中的安全漏洞,指出深度学习路由易受攻击。

Life-Cycle Routing Vulnerabilities of LLM Router

  • 对比白盒/黑盒攻击与后门攻击,评估主流路由模型鲁棒性
  • 基于DNN的路由在训练和推理中均最脆弱,因特征提取能力放大风险
  • 无训练路由因无可操纵参数,对各类攻击防御力最强,适合高安全场景

大型语言模型(LLMs)在自然语言处理中取得显著成功,但其性能与计算成本差异较大。LLM路由器在动态平衡这些权衡中起关键作用。尽管以往研究主要关注路由效率,但整个路由器生命周期(从训练到推理)中的安全漏洞仍基本未被探索。本文全面研究了LLM路由器的全生命周期路由漏洞。我们在多种代表性路由模型上,于广泛实验设置下评估了白盒与黑盒对抗鲁棒性,以及后门鲁棒性。实验发现:1)主流基于DNN的路由器在对抗与后门攻击中表现出最弱鲁棒性,主要因其强特征提取能力在训练和推理阶段均放大了漏洞;2)无训练路由器在不同攻击类型下展现出最强鲁棒性,得益于无可学习参数,难以被操纵。这些发现揭示了覆盖整个生命周期的关键安全风险,并为构建更鲁棒的模型提供了洞见。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable success in natural language processing, yet their performance and computational costs vary significantly. LLM routers play a crucial role in dynamically balancing these trade-offs. While previous studies have primarily focused on routing efficiency, security vulnerabilities throughout the entire LLM router life cycle, from training to inference, remain largely unexplored. In this paper, we present a comprehensive investigation into the life-cycle routing vulnerabilities of LLM routers. We evaluate both white-box and black-box adversarial robustness, as well as backdoor robustness, across several representative routing models under extensive experimental settings. Our experiments uncover several key findings: 1) Mainstream DNN-based routers tend to exhibit the weakest adversarial and backdoor robustness, largely due to their strong feature extraction capabilities that amplify vulnerabilities during both training and inference; 2) Training-free routers demonstrate the strongest robustness across different attack types, benefiting from the absence of learnable parameters that can be manipulated. These findings highlight critical security risks spanning the entire life cycle of LLM routers and provide insights for developing more robust models.

大模型路由安全漏洞对抗攻击后门攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。