用隐空间扰动生成难以察觉的表格数据对抗样本
Crafting Imperceptible On-Manifold Adversarial Attacks for Tabular Data
- 通过混合输入VAE将类别与数值特征统一到隐空间
- 在6个数据集上实现更低离群率和更高分布一致性
- 适合关注表格数据安全与模型鲁棒性的研究者
表格数据的对抗攻击因类别与数值特征混杂而面临独特挑战。传统基于梯度的方法依赖ℓ_p范数约束,常导致偏离原始数据分布。为此,我们提出一种基于混合输入变分自编码器(VAE)的隐空间扰动框架,将类别嵌入与数值特征融合至统一隐流形,实现统计一致性保持的对抗样本生成。引入分布内成功率(IDSR)联合评估攻击效果与分布对齐程度。在六个公开数据集和三种模型架构上的实验表明,该方法显著降低离群率并提升性能一致性,优于传统输入空间攻击及其他从图像领域迁移的VAE方法。超参数敏感性、稀疏性控制及生成架构的综合分析显示,VAE类攻击效果高度依赖重建质量与训练数据量。在满足条件时,其实际效用与稳定性均优于输入空间方法。本工作强调了维持流形内扰动对生成真实且鲁棒对抗样本的重要性。
原文摘要 · Abstract (English)
Adversarial attacks on tabular data present unique challenges due to the heterogeneous nature of mixed categorical and numerical features. Unlike images where pixel perturbations maintain visual similarity, tabular data lacks intuitive similarity metrics, making it difficult to define imperceptible modifications. Additionally, traditional gradient-based methods prioritise $\ell_p$-norm constraints, often producing adversarial examples that deviate from the original data distributions. To address this, we propose a latent-space perturbation framework using a mixed-input Variational Autoencoder (VAE) to generate statistically consistent adversarial examples. The proposed VAE integrates categorical embeddings and numerical features into a unified latent manifold, enabling perturbations that preserve statistical consistency. We introduce In-Distribution Success Rate (IDSR) to jointly evaluate attack effectiveness and distributional alignment. Evaluation across six publicly available datasets and three model architectures demonstrates that our method achieves substantially lower outlier rates and more consistent performance compared to traditional input-space attacks and other VAE-based methods adapted from image domain approaches, achieving substantially lower outlier rates and higher IDSR across six datasets and three model architectures. Our comprehensive analyses of hyperparameter sensitivity, sparsity control, and generative architecture demonstrate that the effectiveness of VAE-based attacks depends strongly on reconstruction quality and the availability of sufficient training data. When these conditions are met, the proposed framework achieves superior practical utility and stability compared with input-space methods. This work underscores the importance of maintaining on-manifold perturbations for generating realistic and robust adversarial examples in tabular domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。