arXiv:2508.02312cs.CRcs.AI2025-08综述被引 2

梳理大模型数据安全风险与防御策略,助力可信AI发展

A Survey on Data Security in Large Language Models

  • 系统归纳大模型训练数据中的安全威胁类型
  • 总结对抗训练、强化学习等主流防御方法
  • 适合研究者与政策制定者参考,推动安全治理

大型语言模型(LLMs)作为自然语言处理的基石,支撑文本生成、机器翻译和对话系统等应用。然而,其依赖海量未经筛选的数据训练,面临严重数据安全风险:有害或恶意数据可能导致模型产生有毒输出、幻觉,或引发提示注入、数据投毒等攻击。随着大模型广泛应用于关键系统,保障数据安全对维护用户信任和系统可靠性至关重要。本文全面综述大模型面临的主要数据安全风险,评述对抗训练、基于人类反馈的强化学习(RLHF)、数据增强等现有防御策略,并分类分析用于评估鲁棒性与安全性的相关数据集。最后,提出安全模型更新、可解释性驱动的防护机制及有效治理框架等未来研究方向,旨在推动大模型技术的安全与负责任发展。本工作面向研究人员、从业者与政策制定者,促进数据安全领域的进展。

原文摘要 · Abstract (English)

Large Language Models (LLMs), now a foundation in advancing natural language processing, power applications such as text generation, machine translation, and conversational systems. Despite their transformative potential, these models inherently rely on massive amounts of training data, often collected from diverse and uncurated sources, which exposes them to serious data security risks. Harmful or malicious data can compromise model behavior, leading to issues such as toxic output, hallucinations, and vulnerabilities to threats such as prompt injection or data poisoning. As LLMs continue to be integrated into critical real-world systems, understanding and addressing these data-centric security risks is imperative to safeguard user trust and system reliability. This survey offers a comprehensive overview of the main data security risks facing LLMs and reviews current defense strategies, including adversarial training, RLHF, and data augmentation. Additionally, we categorize and analyze relevant datasets used for assessing robustness and security across different domains, providing guidance for future research. Finally, we highlight key research directions that focus on secure model updates, explainability-driven defenses, and effective governance frameworks, aiming to promote the safe and responsible development of LLM technology. This work aims to inform researchers, practitioners, and policymakers, driving progress toward data security in LLMs.

大模型安全数据安全防御策略综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。