arXiv:2502.03721cs.CRcs.LG2025-02被引 3

利用语义相似性检测语义通信中的后门攻击,无需改动模型或数据格式。

Detecting Backdoor Attacks via Similarity in Semantic Communication Systems

  • 通过分析语义特征空间的异常偏差来识别中毒样本。
  • 在不同污染比例下检测准确率和召回率均表现优异。
  • 适合关注语义通信安全且希望避免模型修改的研究者。

语义通信系统利用生成式人工智能(GAI)传输语义信息而非原始数据,有望革新现代通信。然而,这类系统易受后门攻击影响——攻击者在训练数据中嵌入恶意触发器,导致中毒样本推理错误,而干净样本不受影响。现有防御方法存在缺陷:如神经元剪枝会降低干净样本的推理性能,或“语义盾”等方法要求图像-文本对的数据格式。为此,本文提出一种不修改模型结构、不强制数据格式的防御机制,基于语义相似性检测后门攻击。通过分析语义特征空间的偏差并建立阈值检测框架,有效识别中毒样本。实验结果表明,在不同污染比例下,该方法均具备高检测准确率与召回率,验证了其显著有效性。

原文摘要 · Abstract (English)

Semantic communication systems, which leverage Generative AI (GAI) to transmit semantic meaning rather than raw data, are poised to revolutionize modern communications. However, they are vulnerable to backdoor attacks, a type of poisoning manipulation that embeds malicious triggers into training datasets. As a result, Backdoor attacks mislead the inference for poisoned samples while clean samples remain unaffected. The existing defenses may alter the model structure (such as neuron pruning that potentially degrades inference performance on clean inputs, or impose strict requirements on data formats (such as ``Semantic Shield" that requires image-text pairs). To address these limitations, this work proposes a defense mechanism that leverages semantic similarity to detect backdoor attacks without modifying the model structure or imposing data format constraints. By analyzing deviations in semantic feature space and establishing a threshold-based detection framework, the proposed approach effectively identifies poisoned samples. The experimental results demonstrate high detection accuracy and recall across varying poisoning ratios, underlining the significant effectiveness of our proposed solution.

语义通信后门攻击生成式AI安全检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。