arXiv:2503.08644cs.CLcs.AI2025-03ACL被引 7

指令跟随检索器可能被恶意利用,泄露有害信息。

Exploiting Instruction-Following Retrievers for Malicious Information Retrieval

  • 测试六款主流检索器,发现超50%恶意查询能匹配到有害内容。
  • LLM2Vec在恶意查询中正确召回率达61.35%。
  • 即使安全对齐的Llama3模型,也可能因检索到的有害内容而生成危险输出。

指令跟随型检索器在实际应用中广泛与大语言模型结合使用,但其日益增强的搜索能力带来的安全风险尚未得到充分研究。我们实证研究了检索器在直接使用及检索增强生成场景下满足恶意查询的能力。具体而言,评估了包括NV-Embed和LLM2Vec在内的六款领先检索器,发现对于超过50%的恶意请求,多数检索器能选择相关有害段落。例如,LLM2Vec在我们的恶意查询中正确选择比例达61.35%。我们进一步揭示了一种新兴风险:攻击者可通过操纵指令,利用检索器的指令遵循能力主动暴露高度相关的有害信息。最后,我们证明,即便经过安全对齐的LLM(如Llama3),在上下文引入有害检索结果后仍可能生成恶意内容。综上,研究结果凸显了检索器能力提升所带来的恶意滥用风险。

原文摘要 · Abstract (English)

Instruction-following retrievers have been widely adopted alongside LLMs in real-world applications, but little work has investigated the safety risks surrounding their increasing search capabilities. We empirically study the ability of retrievers to satisfy malicious queries, both when used directly and when used in a retrieval augmented generation-based setup. Concretely, we investigate six leading retrievers, including NV-Embed and LLM2Vec, and find that given malicious requests, most retrievers can (for >50% of queries) select relevant harmful passages. For example, LLM2Vec correctly selects passages for 61.35% of our malicious queries. We further uncover an emerging risk with instruction-following retrievers, where highly relevant harmful information can be surfaced by exploiting their instruction-following capabilities. Finally, we show that even safety-aligned LLMs, such as Llama3, can satisfy malicious requests when provided with harmful retrieved passages in-context. In summary, our findings underscore the malicious misuse risks associated with increasing retriever capability.

检索安全恶意查询LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。