arXiv:2511.07888cs.CL2025-11AAAI

通过净化嵌入流形,同时提升文本分类的鲁棒性和准确率。

Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold Purification

  • 用流形学习识别并修正对抗样本的嵌入表示
  • 在三个数据集上实现更强鲁棒性且干净数据性能不降反升
  • 适合关注模型安全与精度兼得的研究者

文本分类中提升模型对抗攻击的鲁棒性通常会降低其在干净数据上的性能。我们提出,可通过建模编码器嵌入空间中清洁样本的分布来解决这一问题。为此,我们设计了基于流形校正的因果流(MC²F),该系统直接作用于句子嵌入。其中分层黎曼连续归一化流(SR-CNF)学习清洁数据流形的密度分布,并识别出分布外的嵌入;随后由测地线净化求解器将其沿最短路径投影回学习到的流形上,恢复语义一致的干净表示。我们在三个数据集和多种对抗攻击下进行了全面评估。结果表明,所提方法不仅在对抗鲁棒性上达到新基准,且对干净数据的性能完全保持甚至略有提升。

原文摘要 · Abstract (English)

A persistent challenge in text classification (TC) is that enhancing model robustness against adversarial attacks typically degrades performance on clean data. We argue that this challenge can be resolved by modeling the distribution of clean samples in the encoder embedding manifold. To this end, we propose the Manifold-Correcting Causal Flow (MC^2F), a two-module system that operates directly on sentence embeddings. A Stratified Riemannian Continuous Normalizing Flow (SR-CNF) learns the density of the clean data manifold. It identifies out-of-distribution embeddings, which are then corrected by a Geodesic Purification Solver. This solver projects adversarial points back onto the learned manifold via the shortest path, restoring a clean, semantically coherent representation. We conducted extensive evaluations on text classification (TC) across three datasets and multiple adversarial attacks. The results demonstrate that our method, MC^2F, not only establishes a new state-of-the-art in adversarial robustness but also fully preserves performance on clean data, even yielding modest gains in accuracy.

文本分类对抗鲁棒性嵌入流形净化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。