arXiv:2607.13801cs.CRcs.AI2026-07

为大模型网络入侵检测设计抗流量操纵的强鲁棒防御方法

Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection

  • 仅在攻击者可操控的特征子空间注入高斯噪声,实现精准防御
  • 在多个数据集上认证准确率提升至55%-100%,最大认证半径超基准值5倍
  • 适合关注大模型安全性的研究人员与工业级入侵检测系统开发者

基于大语言模型(LLM)的入侵检测系统(IDS)日益受到关注,但其对可行流量操纵的鲁棒性仍主要依赖经验验证。本文提出流量感知随机平滑(TA-RS),一种无需依赖具体分类器的认证防御机制,在微调和认证阶段仅向远程攻击者可修改的直接可控(DC)子空间注入高斯噪声,使平滑分布与攻击者可控空间对齐。实验发现,标准随机平滑在三个(模型,数据集)组合中认证准确率仅为14%-33%(低于随机水平),仅第四组达57%(较噪声增强结果低43个百分点);而噪声增强微调后,两个基准数据集上的认证准确率恢复至68%-100%(sigma=0.25)。在L_inf等效阈值R_inf = epsilon×√|DC|(epsilon=0.05)下,TA-RS在CIC-IDS-2018和HIKARI-2021上实现55%-100%认证准确率,中位认证半径R≈0.45-0.96,超出R_inf 1.8-5倍(sigma=0.25-1.00)。相比同训练配方的等距随机平滑基线,其优势在不同数据集间差异显著(在CIC-IDS-2018上高出4-19个百分点);而与无方向性随机平滑基线相比,差距高达72个百分点,主要源于训练-认证不一致:无方向性测试噪声扰动了攻击者无法利用的不可控特征,导致弃权率高达68%。在RT-IoT2022上,原微调方案失效,但通过增强噪声注入,仍可恢复至76%/69%认证准确率(LLaMA3-8B/Qwen3-8B)。

原文摘要 · Abstract (English)

Large language model (LLM)-based intrusion detection systems (IDS) are increasingly studied for security monitoring, yet their robustness against feasible traffic manipulation remains largely empirical. We present Traffic-Aware Randomized Smoothing (TA-RS), a classifier-agnostic certified defense that injects Gaussian noise exclusively into the directly controllable (DC) subspace -- features a remote attacker can modify -- during both fine-tuning and certification, aligning the smoothing distribution with the attacker-controllable subspace. We identify a critical prerequisite: applying standard randomized smoothing to clean-trained LLM-IDS yields weak certified accuracy in three of four (model, dataset) pairs tested (14-33%, at or below random) and only 57% in the fourth (43 pp below the noise-augmented result); noise-augmented fine-tuning recovers to 68-100% on two of three benchmark datasets (at sigma=0.25). At the L_inf-equivalent threshold R_inf = epsilon*sqrt(|DC|) (epsilon=0.05), TA-RS achieves 55-100% certified accuracy on CIC-IDS-2018 and HIKARI-2021, with median certified radii (R approx 0.45-0.96) exceeding R_inf by 1.8-5x (across sigma=0.25-1.00). Against a fairly trained iso-trained RS baseline the residual advantage is dataset-dependent (4-19 pp on CIC-IDS-2018). The larger gap -- up to 72 pp against an isotropic RS baseline that shares the DC-noise-augmented training recipe -- primarily reflects the training-certification mismatch rather than DC alignment alone: isotropic test-time noise perturbs uncontrollable features the attacker cannot exploit, triggering abstention rates up to 68%. RT-IoT2022 probes the limits of the method: it fails under the default fine-tuning recipe but recovers to 76%/69% certified accuracy (LLaMA3-8B/Qwen3-8B) when noise augmentation is increased.

大模型安全入侵检测鲁棒性防御随机平滑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。