系统梳理大模型责任风险与防护策略,覆盖全生命周期
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
- 按开发使用四阶段构建统一治理框架
- 针对幻觉、隐私泄露等风险提出针对性防御方案
- 适合关注AI安全与合规的开发者和研究者
大型语言模型(LLMs)在支持现实应用和带来社会价值方面具有巨大潜力,但仍面临隐私泄露、幻觉输出和价值观错位等固有风险,且可能被恶意利用生成有害内容或用于不道德目的。为此,本文系统综述了近期旨在缓解这些问题的进展,涵盖大模型发展的四个阶段:数据收集与预训练、微调与对齐、提示与推理、后处理与审计。重点阐述了在隐私保护、幻觉减少、价值观对齐、毒性消除及越狱防御等方面的最新进展。相较于以往仅聚焦单一维度的综述,本工作提出一个统一框架,整合多种治理维度,为提升大模型在真实场景中的可靠性提供全面视角。
原文摘要 · Abstract (English)
While large language models (LLMs) present significant potential for supporting numerous real-world applications and delivering positive social impacts, they still face significant challenges in terms of the inherent risk of privacy leakage, hallucinated outputs, and value misalignment, and can be maliciously used for generating toxic content and unethical purposes after been jailbroken. Therefore, in this survey, we present a comprehensive review of recent advancements aimed at mitigating these issues, organized across the four phases of LLM development and usage: data collecting and pre-training, fine-tuning and alignment, prompting and reasoning, and post-processing and auditing. We elaborate on the recent advances for enhancing the performance of LLMs in terms of privacy protection, hallucination reduction, value alignment, toxicity elimination, and jailbreak defenses. In contrast to previous surveys that focus on a single dimension of responsible LLMs, this survey presents a unified framework that encompasses these diverse dimensions, providing a comprehensive view of enhancing LLMs to better serve real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。