arXiv:2606.16821cs.CLcs.CR2026-06被引 2

测试大模型搜索代理被网页内容操纵时的推荐可靠性,发现不同模型漏洞差异显著。

How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation

论文配图:How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation
图 1 · 摘自论文原文
  • 构建可控框架SearchGEO,模拟五类攻击并评估推荐结果被篡改程度。
  • 攻击成功率在0%到31.4%之间,不同模型对同一攻击响应差异大。
  • 揭示模型在信任与拒绝行为上的极端倾向,适合关注安全性的研究者。

基于大语言模型(LLM)的搜索代理会将开放网络内容合成用户可执行的建议,存在攻击者发布的页面被转化为可信主张的风险。我们提出SearchGEO,一个用于测量LLM搜索代理推荐腐败的受控评估框架,包含网页证据操控流程、五种攻击模式分类及多维度输出指标。在13个LLM后端上对308个案例进行评估。结果显示,各后端漏洞模式各异:整体攻击成功率(ASR)从Claude-Sonnet-4.6的0.0%到Gemini-3-Flash的31.4%不等,最强攻击方式因模型家族而异,相同部署结构可能在不同后端上放大或降低ASR。通过辅助代理技能探测任务(将推荐转为安装命令),暴露了原本稳健模型间的明显分歧:Claude过度拒绝,GPT过度信任。这些发现表明,应将对抗性搜索内容下的推荐可靠性视为后端安全性评估的核心维度。

原文摘要 · Abstract (English)

Large language model (LLM)-based search agents synthesize open-web content into actionable recommendations on behalf of users, creating a risk that attacker-published pages are transformed into endorsed claims. We introduce SearchGEO, a controlled evaluation framework for measuring endorsement corruption in LLM-based web-search agents, combining a web-evidence manipulation pipeline, a five-mode attack taxonomy, and multiple output-level metrics. We evaluate 13 LLM backends on 308 cases each. Results show that vulnerability patterns vary across backends: overall attack success rate (ASR) ranges from 0.0% on Claude-Sonnet-4.6 to 31.4% on Gemini-3-Flash, the strongest attack mode differs by model family, and the same deployment scaffold could amplify or decrease ASR on different backends. An auxiliary agent-skill probe, where endorsement becomes an install command, exposes a sharp split among otherwise robust backends: Claude over-rejects while GPT over-trusts. These findings argue for treating recommendation reliability under adversarial search content as a first-class dimension of backend safety evaluation.

大模型安全搜索代理信任风险对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。