arXiv:2410.20315cs.CLcs.AI2024-10被引 2

测试密集检索模型对分词器攻击的脆弱性,发现监督模型易受影响,无监督模型更抗干扰。

Deep Learning Based Dense Retrieval: A Comparative Study

  • 通过污染分词器测试BERT、DPR等模型的鲁棒性
  • 小扰动即可导致监督模型检索准确率大幅下降
  • 适合关注检索系统安全性的研究人员参考

密集检索器在多种信息检索任务中表现优异,但其对分词器污染的鲁棒性尚未深入研究。本文评估了BERT、Dense Passage Retrieval(DPR)、Contriever、SimCSE和ANCE等模型在分词器被攻击时的表现。结果表明,监督模型如BERT和DPR在分词器受损时性能显著下降,而无监督模型如ANCE表现出更强的韧性。实验显示,微小扰动即可严重损害检索精度,凸显了在关键应用中构建鲁棒防御机制的必要性。

原文摘要 · Abstract (English)

Dense retrievers have achieved state-of-the-art performance in various information retrieval tasks, but their robustness against tokenizer poisoning remains underexplored. In this work, we assess the vulnerability of dense retrieval systems to poisoned tokenizers by evaluating models such as BERT, Dense Passage Retrieval (DPR), Contriever, SimCSE, and ANCE. We find that supervised models like BERT and DPR experience significant performance degradation when tokenizers are compromised, while unsupervised models like ANCE show greater resilience. Our experiments reveal that even small perturbations can severely impact retrieval accuracy, highlighting the need for robust defenses in critical applications.

密集检索模型安全分词器攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。