用知识图谱识别诈骗关键词,提升大模型抗骗能力。
FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks
- 构建诈骗手法-关键词知识图谱,关联可疑文本与欺诈技术。
- 在5类诈骗攻击中,对4个主流大模型均显著优于现有防御方法。
- 提供可解释的提示线索,适合安全敏感场景使用。
大语言模型广泛应用于合同审查、求职申请等关键自动化流程,但易受欺诈信息干扰,导致严重后果。现有防御方法在有效性、可解释性和泛化性方面存在局限。为此,我们提出FraudShield框架,通过分析欺诈策略构建并优化欺诈手法-关键词知识图谱,捕捉可疑文本与欺诈技术之间的高置信度关联。该结构化知识图谱通过标注关键词并提供佐证,增强输入信息,引导大模型生成更安全的响应。大量实验表明,FraudShield在四个主流大模型上对五类典型诈骗攻击均持续优于现有最优防御方案,并能提供可解释的生成依据。
原文摘要 · Abstract (English)
Large language models (LLMs) have been widely integrated into critical automated workflows, including contract review and job application processes. However, LLMs are susceptible to manipulation by fraudulent information, which can lead to harmful outcomes. Although advanced defense methods have been developed to address this issue, they often exhibit limitations in effectiveness, interpretability, and generalizability, particularly when applied to LLM-based applications. To address these challenges, we introduce FraudShield, a novel framework designed to protect LLMs from fraudulent content by leveraging a comprehensive analysis of fraud tactics. Specifically, FraudShield constructs and refines a fraud tactic-keyword knowledge graph to capture high-confidence associations between suspicious text and fraud techniques. The structured knowledge graph augments the original input by highlighting keywords and providing supporting evidence, guiding the LLM toward more secure responses. Extensive experiments show that FraudShield consistently outperforms state-of-the-art defenses across four mainstream LLMs and five representative fraud types, while also offering interpretable clues for the model's generations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。