arXiv:2602.01369cs.CV2026-02被引 3

揭示视频MoE模型的组件级脆弱性并提出协同防御方法

Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts

  • 针对路由器和专家模块分别设计攻击,发现独立弱点
  • 联合扰动路由器与专家,暴露架构协同漏洞,攻击效果提升显著
  • 提出协同对抗训练法,兼顾防御独立与协同弱点,效率高

Mixture-of-Experts(MoE)在视频理解任务中表现优异,但其对抗鲁棒性尚未充分研究。现有攻击方法常将MoE视为整体结构,忽视了路由器与专家模块的独立与协同弱点。为此,本文提出时序Lipschitz引导攻击(TLGA),首先针对路由器设计攻击,揭示其独立弱点;在此基础上,提出联合时序Lipschitz引导攻击(J-TLGA),协同扰动路由器与专家,显著放大对抗效应,暴露MoE架构的协同弱点(即“阿喀琉斯之踵”)。基于此,进一步提出联合时序Lipschitz对抗训练(J-TLAT),通过联合训练增强组件级鲁棒性。该框架可即插即用,推理成本比密集模型降低60%以上,在多种数据集与架构上持续提升对抗鲁棒性,有效缓解MoE的独立与协同弱点。

原文摘要 · Abstract (English)

Mixture-of-Experts (MoE) has demonstrated strong performance in video understanding tasks, yet its adversarial robustness remains underexplored. Existing attack methods often treat MoE as a unified architecture, overlooking the independent and collaborative weaknesses of key components such as routers and expert modules. To fill this gap, we propose Temporal Lipschitz-Guided Attacks (TLGA) to thoroughly investigate component-level vulnerabilities in video MoE models. We first design attacks on the router, revealing its independent weaknesses. Building on this, we introduce Joint Temporal Lipschitz-Guided Attacks (J-TLGA), which collaboratively perturb both routers and experts. This joint attack significantly amplifies adversarial effects and exposes the Achilles' Heel (collaborative weaknesses) of the MoE architecture. Based on these insights, we further propose Joint Temporal Lipschitz Adversarial Training (J-TLAT). J-TLAT performs joint training to further defend against collaborative weaknesses, enhancing component-wise robustness. Our framework is plug-and-play and reduces inference cost by more than 60% compared with dense models. It consistently enhances adversarial robustness across diverse datasets and architectures, effectively mitigating both the independent and collaborative weaknesses of MoE.

视频理解MoE对抗攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。