检测并清除开放无线网络中深度强化学习组件的隐蔽后门攻击
ORAN-DEFEND: Subspace Detection and Sanitization of Backdoor DRL xApps in Open RAN
- 通过奇异值分解构建安全子空间,对关键性能指标进行投影净化
- 在真实数据集上实现100%性能恢复和99.5%以上防御成功率
- 揭示线性投影防御的固有极限,适用于网络安全研究人员
开放无线接入网(O-RAN)越来越多地将近实时控制权交给第三方提供的深度强化学习(DRL)xApps,形成新的供应链攻击面。后门策略在正常情况下表现良好,但一旦攻击者向关键性能指标(KPI)遥测注入隐蔽触发信号,便会执行损害服务质量(QoS)的恶意操作。我们提出ORAN-DEFEND,一种无需重新训练的封装式防御机制,通过奇异值分解(SVD)从少量可信干净回放数据中估计安全子空间,并将每个KPI窗口投影至该子空间以实现净化。我们从理论与实验两方面证明了精确的恢复条件:当触发信号能量集中在安全子空间的正交补空间时,防御有效,并以触发信号的$\ ext{E}_\perp$能量占比量化该边界。在Colosseum COLORAN数据集上,我们评估了四种结构不同的DRL后门攻击(包括TrojDRL、SleeperNets、BadRL和Q-Incept),覆盖内环与外环中毒场景,验证了在满足子空间假设条件下,所有攻击均实现100%性能恢复与≥99.5%的防御成功率。几何消融实验揭示任意线性投影防御的内在局限:当触发信号与合法信号共位时,$ ext{E}_\perp$能量占比决定恢复程度单调变化,此时线性残差检测器退化为随机猜测,而非线性分类器仍能保持完全可分性。
原文摘要 · Abstract (English)
Open Radio Access Networks (O-RAN) increasingly delegate near-real-time control to deep reinforcement learning (DRL) xApps obtained from third-party vendors, creating a new supply-chain attack surface. A backdoor policy behaves optimally until an adversary injects a covert trigger into the observed key performance indicator (KPI) telemetry, at which point it issues harmful control actions that degrade quality of service (QoS). We present ORAN-DEFEND, a retraining-free wrapper that sanitizes a frozen, potentially compromised xApp by projecting each KPI window onto a safe subspace estimated from a small number of trusted clean rollouts via singular value decomposition (SVD). We establish, both analytically and empirically, a precise recovery condition: the defense succeeds if the trigger energy concentrates in the orthogonal complement of the safe subspace, and we quantify this boundary through the trigger's $\Eperp$ energy fraction. On the Colosseum COLORAN dataset, we evaluate four structurally distinct DRL backdoor attacks, like TrojDRL, SleeperNets, BadRL, and Q-Incept, spanning inner-loop and outer-loop poisoning regimes and demonstrate $100\%$ return recovery and $\geq99.5\%$ defense success rate across all four when the subspace assumption holds. A geometry ablation reveals an intrinsic and previously uncharacterized limit of any linear projection defense: when the trigger collocates with the legitimate signal, the $\Eperp$ energy fraction governs recovery monotonically, and the linear residual detector collapses to chance even while a nonlinear classifier retains perfect separability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。