arXiv:2502.00384cs.CRcs.LG2025-02NeurIPS

用可解释性技术破解深度学习侧信道攻击的黑箱,揭示其如何利用物理信号窃密。

Interpreting Emergent Features in Deep Learning-based Side-channel Analysis

  • 通过机制可解释性分析模型在侧信道数据中捕捉到的泄漏特征。
  • 在输入稀疏、准确率低的情况下仍能还原出秘密掩码。
  • 帮助安全评估者从黑盒攻击转向白盒防御,适合硬件安全研究者。

侧信道分析(SCA)通过利用设备无意泄露的物理信号,从安全设备中提取秘密信息,也被用于安全认证。近年来,深度学习在SCA中表现卓越,但缺乏可解释性。本文采用机制可解释性方法,分析训练用于SCA的神经网络,揭示模型如何利用侧信道迹线中的特定泄漏。聚焦性能突增现象,反向解析模型学到的表示,最终恢复出秘密掩码,推动评估从黑盒走向白盒。结果表明,该方法可在真实SCA场景中应用,即使有效输入稀疏、模型精度较低,且侧信道防护阻止常规输入干预。

原文摘要 · Abstract (English)

Side-channel analysis (SCA) poses a real-world threat by exploiting unintentional physical signals to extract secret information from secure devices. Evaluation labs also use the same techniques to certify device security. In recent years, deep learning has emerged as a prominent method for SCA, achieving state-of-the-art attack performance at the cost of interpretability. Understanding how neural networks extract secrets is crucial for security evaluators aiming to defend against such attacks, as only by understanding the attack can one propose better countermeasures. In this work, we apply mechanistic interpretability to neural networks trained for SCA, revealing \textit{how} models exploit \textit{what} leakage in side-channel traces. We focus on sudden jumps in performance to reverse engineer learned representations, ultimately recovering secret masks and moving the evaluation process from black-box to white-box. Our results show that mechanistic interpretability can scale to realistic SCA settings, even when relevant inputs are sparse, model accuracies are low, and side-channel protections prevent standard input interventions.

侧信道攻击可解释性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。