通过扰动嵌入实现对排序模型的隐蔽攻击,96%目标文档可被提升至前10名。
EMPRA: Embedding Perturbation Rank Attack against Neural Ranking Models
- 基于句子级嵌入扰动,生成人类难以察觉的对抗文本。
- 在MS MARCO数据集上,96%原排名51-100的目标文档被重排至前10。
- 无需替代模型,适用于真实场景下的黑盒攻击,适用性强。
近期研究显示,神经信息检索技术可能面临对抗攻击。对抗攻击旨在操纵文档排序,使用户接触到特定内容。本文提出嵌入扰动排序攻击(EMPRA),一种针对黑盒神经排序模型(NRMs)的新方法。EMPRA通过扰动句子级嵌入,引导其趋向与查询相关的语境,同时保持语义完整性,生成能无缝融入原文且人眼难以察觉的对抗文本。我们在广泛使用的MS MARCO V1段落集合上进行了全面评估,结果表明EMPRA在促进特定目标文档在排名结果中表现方面,对多种先进基线方法均有效。具体而言,几乎96%原本排名51-100的目标文档被成功重排至前10位。此外,EMPRA不依赖替代模型生成对抗文本,增强了其在真实环境中的鲁棒性。
原文摘要 · Abstract (English)
Recent research has shown that neural information retrieval techniques may be susceptible to adversarial attacks. Adversarial attacks seek to manipulate the ranking of documents, with the intention of exposing users to targeted content. In this paper, we introduce the Embedding Perturbation Rank Attack (EMPRA) method, a novel approach designed to perform adversarial attacks on black-box Neural Ranking Models (NRMs). EMPRA manipulates sentence-level embeddings, guiding them towards pertinent context related to the query while preserving semantic integrity. This process generates adversarial texts that seamlessly integrate with the original content and remain imperceptible to humans. Our extensive evaluation conducted on the widely-used MS MARCO V1 passage collection demonstrate the effectiveness of EMPRA against a wide range of state-of-the-art baselines in promoting a specific set of target documents within a given ranked results. Specifically, EMPRA successfully achieves a re-ranking of almost 96% of target documents originally ranked between 51-100 to rank within the top 10. Furthermore, EMPRA does not depend on surrogate models for adversarial text generation, enhancing its robustness against different NRMs in realistic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。