系统梳理大模型代理的安全、隐私与伦理风险,提出新分类框架。
Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents
- 按风险来源与影响构建新型分类框架,覆盖跨模块跨阶段威胁。
- 归纳六类核心特征,分析现有研究的局限性与实践案例风险。
- 面向研究人员和政策制定者,指明数据、方法与治理未来方向。
随着大语言模型(LLMs)的发展,基于Transformer的模型在自然语言处理(NLP)任务中取得突破性进展,催生了一系列以LLM为核心控制单元的智能体。尽管LLM在各类任务中表现优异,其面临的安全与隐私威胁在代理场景下更加严峻。为提升基于LLM应用的可靠性,已有大量研究从不同角度评估并缓解这些风险。为帮助研究者全面理解各类风险,本文系统收集并分析了此类智能体所面临的威胁。针对以往分类体系难以处理跨模块、跨阶段威胁的问题,本文提出一种基于风险源与影响的新分类框架。基于六个关键特征,总结当前研究进展并分析其局限性。随后选取四个代表性智能体进行案例研究,剖析其实际应用中的潜在风险。最后,从数据、方法与政策三个维度,提出未来研究方向。
原文摘要 · Abstract (English)
With the continuous development of large language models (LLMs), transformer-based models have made groundbreaking advances in numerous natural language processing (NLP) tasks, leading to the emergence of a series of agents that use LLMs as their control hub. While LLMs have achieved success in various tasks, they face numerous security and privacy threats, which become even more severe in the agent scenarios. To enhance the reliability of LLM-based applications, a range of research has emerged to assess and mitigate these risks from different perspectives. To help researchers gain a comprehensive understanding of various risks, this survey collects and analyzes the different threats faced by these agents. To address the challenges posed by previous taxonomies in handling cross-module and cross-stage threats, we propose a novel taxonomy framework based on the sources and impacts. Additionally, we identify six key features of LLM-based agents, based on which we summarize the current research progress and analyze their limitations. Subsequently, we select four representative agents as case studies to analyze the risks they may face in practical use. Finally, based on the aforementioned analyses, we propose future research directions from the perspectives of data, methodology, and policy, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。