arXiv:2504.05804cs.IRcs.AI2025-04被引 19

用隐蔽提示词操控大模型排序,不被发现还能有效提升目标排名

StealthRank: LLM Ranking Manipulation via Stealthy Prompt Optimization

  • 基于能量优化与朗之万动力学生成隐蔽对抗提示
  • 在多个大模型上成功提升目标项排名且无明显异常痕迹
  • 适合研究模型安全或想了解对抗攻击的读者

将大语言模型(LLMs)集成到信息检索系统中带来了新的攻击面,尤其是针对对抗性排序操纵。我们提出了一种名为 StealthRank 的新型对抗攻击方法,能够在保持文本流畅性的同时,隐秘地操纵由大模型驱动的排序系统。与以往常引入可检测异常的方法不同,StealthRank 采用基于能量的优化框架结合朗之万动力学,生成嵌入在项目或文档描述中的对抗性文本序列(称为 SRPs),以微妙但有效的方式影响大模型的排序机制。我们在多个 LLM 上评估了 StealthRank,结果表明其能够隐蔽地提升目标项目的排名,同时避免显式的操纵痕迹。实验显示,StealthRank 在有效性与隐蔽性方面均持续优于现有最先进对抗排序基线,揭示了大模型驱动排序系统的严重安全隐患。代码已公开于 https://github.com/Tangyiming205069/controllable-seo。

原文摘要 · Abstract (English)

The integration of large language models (LLMs) into information retrieval systems introduces new attack surfaces, particularly for adversarial ranking manipulations. We present $\textbf{StealthRank}$, a novel adversarial attack method that manipulates LLM-driven ranking systems while maintaining textual fluency and stealth. Unlike existing methods that often introduce detectable anomalies, StealthRank employs an energy-based optimization framework combined with Langevin dynamics to generate StealthRank Prompts (SRPs)-adversarial text sequences embedded within item or document descriptions that subtly yet effectively influence LLM ranking mechanisms. We evaluate StealthRank across multiple LLMs, demonstrating its ability to covertly boost the ranking of target items while avoiding explicit manipulation traces. Our results show that StealthRank consistently outperforms state-of-the-art adversarial ranking baselines in both effectiveness and stealth, highlighting critical vulnerabilities in LLM-driven ranking systems. Our code is publicly available at $\href{https://github.com/Tangyiming205069/controllable-seo}{here}$.

大模型安全对抗攻击排序操纵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。