梳理大模型在安全与信任领域的应用挑战与风险
Trust & Safety of LLMs and LLMs in Trust & Safety
- 系统综述大模型在安全领域的应用现状
- 揭示提示注入和越狱攻击等新兴风险
- 适合关注AI伦理与安全的研究者参考
近年来,大型语言模型(LLMs)在自然语言处理任务中展现出卓越能力,但其广泛应用引发了对信任与安全的担忧。本文系统性综述了当前关于大模型在信任与安全领域研究的进展,特别关注大模型在该领域自身的应用。深入探讨了在需保障信任与安全的关键场景中使用大模型所面临的复杂性,整合多篇研究的发现,识别出关键挑战与潜在解决方案,旨在为研究人员和从业者提供对大模型与信任安全之间复杂互动的理解。本综述还提供了在信任与安全中使用大模型的最佳实践建议,并探索了如提示注入和越狱攻击等新兴风险。最终,本研究有助于深化对如何有效且负责任地利用大模型以增强数字领域信任与安全的认识。
原文摘要 · Abstract (English)
In recent years, Large Language Models (LLMs) have garnered considerable attention for their remarkable abilities in natural language processing tasks. However, their widespread adoption has raised concerns pertaining to trust and safety. This systematic review investigates the current research landscape on trust and safety in LLMs, with a particular focus on the novel application of LLMs within the field of Trust and Safety itself. We delve into the complexities of utilizing LLMs in domains where maintaining trust and safety is paramount, offering a consolidated perspective on this emerging trend.\ By synthesizing findings from various studies, we identify key challenges and potential solutions, aiming to benefit researchers and practitioners seeking to understand the nuanced interplay between LLMs and Trust and Safety. This review provides insights on best practices for using LLMs in Trust and Safety, and explores emerging risks such as prompt injection and jailbreak attacks. Ultimately, this study contributes to a deeper understanding of how LLMs can be effectively and responsibly utilized to enhance trust and safety in the digital realm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。