arXiv:2505.17085cs.CRcs.AI2025-05被引 2

通过融合多维弱信号,精准识别社交平台中的隐蔽语言攻击

GSDFuse: Capturing Cognitive Inconsistencies from Multi-Dimensional Weak Signals in Social Media Steganalysis

  • 构建分层多模态特征,捕捉文本碎片化与对话结构中的认知不一致
  • 在极端稀疏场景下仍实现领先检测效果,准确率显著优于现有方法
  • 适合安全分析、网络内容审核等需识别隐蔽信息的场景

社交媒体平台的普及助长了恶意语言隐写行为,带来重大安全风险。传统隐写分析面临两大挑战:一是难以识别由文本碎片化和复杂对话结构引发的细微认知不一致;二是难以稳健聚合多维度弱信号,尤其在极低隐写密度和高复杂度隐写技术下更为困难。此外,数据严重不平衡进一步加剧检测难度。本文提出GSDFuse,一种系统性解决方案,通过分层多模态特征工程捕获多样化信号,采用策略性数据增强缓解稀疏性,设计自适应证据融合机制智能聚合弱信号,并结合判别性嵌入学习提升对细微不一致的敏感度。在多个社交媒体数据集上的实验表明,GSDFuse在复杂对话环境中对高级隐写手法的检测性能达到当前最优水平。源代码已公开于https://github.com/NebulaEmmaZh/GSDFuse。

原文摘要 · Abstract (English)

The ubiquity of social media platforms facilitates malicious linguistic steganography, posing significant security risks. Steganalysis is profoundly hindered by the challenge of identifying subtle cognitive inconsistencies arising from textual fragmentation and complex dialogue structures, and the difficulty in achieving robust aggregation of multi-dimensional weak signals, especially given extreme steganographic sparsity and sophisticated steganography. These core detection difficulties are compounded by significant data imbalance. This paper introduces GSDFuse, a novel method designed to systematically overcome these obstacles. GSDFuse employs a holistic approach, synergistically integrating hierarchical multi-modal feature engineering to capture diverse signals, strategic data augmentation to address sparsity, adaptive evidence fusion to intelligently aggregate weak signals, and discriminative embedding learning to enhance sensitivity to subtle inconsistencies. Experiments on social media datasets demonstrate GSDFuse's state-of-the-art (SOTA) performance in identifying sophisticated steganography within complex dialogue environments. The source code for GSDFuse is available at https://github.com/NebulaEmmaZh/GSDFuse.

隐写分析多模态弱信号社交安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。