AI搜索常泄露恶意链接,研究首次量化风险并提出轻量防御方案。
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search
- 构建威胁模型,系统评估7个主流AI搜索的漏洞
- 正常查询下仍37%响应含恶意网址,直接输入网址风险翻倍
- 用GPT-4.1+URL检测器可降险90%,信息损失仅10.7%
大语言模型(LLM)显著提升了人工智能搜索引擎(AIPSEs)的能力,通过整合外部数据库与已有知识实现精准高效响应。然而,我们发现这些AIPSEs存在引用恶意内容或提供恶意网站链接的风险,导致有害或未经验证信息传播。本研究首次对7个生产级AIPSEs进行安全风险量化分析,系统定义威胁模型、风险类型,并评估其在多种查询下的表现。基于PhishTank、ThreatBook和LevelBlue数据,结果表明:即使使用良性关键词查询,AIPSEs仍频繁生成包含恶意URL的有害内容。直接输入网址会使高风险响应数量增加,而自然语言查询则略有缓解作用。相比传统搜索引擎,AIPSEs在实用性和安全性上均更优。通过在线文档伪造与钓鱼攻击两个案例研究,揭示了现实场景中欺骗AIPSEs的易行性。为缓解风险,我们开发了一种基于代理的防御机制,包含GPT-4.1驱动的内容精炼工具与URL检测器。评估显示,该防御可有效降低风险,仅造成约10.7%可用信息损失。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have significantly enhanced the capabilities of AI-Powered Search Engines (AIPSEs), offering precise and efficient responses by integrating external databases with pre-existing knowledge. However, we observe that these AIPSEs raise risks such as quoting malicious content or citing malicious websites, leading to harmful or unverified information dissemination. In this study, we conduct the first safety risk quantification on seven production AIPSEs by systematically defining the threat model, risk type, and evaluating responses to various query types. With data collected from PhishTank, ThreatBook, and LevelBlue, our findings reveal that AIPSEs frequently generate harmful content that contains malicious URLs even with benign queries (e.g., with benign keywords). We also observe that directly querying a URL will increase the number of main risk-inclusive responses, while querying with natural language will slightly mitigate such risk. Compared to traditional search engines, AIPSEs outperform in both utility and safety. We further perform two case studies on online document spoofing and phishing to show the ease of deceiving AIPSEs in the real-world setting. To mitigate these risks, we develop an agent-based defense with a GPT-4.1-based content refinement tool and a URL detector. Our evaluation shows that our defense can effectively reduce the risk, with only a minor cost of reducing available information by approximately 10.7%. Our research highlights the urgent need for robust safety measures in AIPSEs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。