综述AI在安全、偏见和隐私三方面的可信性挑战与应对思路
Trustworthy AI: Safety, Bias, and Privacy -- A Survey
- 从大模型安全对齐、虚假偏见识别到神经网络成员推断攻击,系统分析可信风险
- 提出针对有毒内容生成、误导性偏见和隐私泄露的检测与缓解框架
- 适合关注AI伦理、安全评估与可解释性的研究人员参考
人工智能系统能力快速提升,但仍面临失效模式、漏洞和偏见等问题。本文综述当前可信人工智能研究进展,聚焦安全、偏见和隐私三大核心挑战。针对安全问题,探讨大语言模型中的安全对齐机制,以防止有害或有毒内容生成;针对偏见问题,重点分析可能误导模型的虚假偏见;针对隐私问题,讨论深度神经网络中的成员推断攻击。文中观点基于作者实验与观察,提供多维度可信性评估视角。
原文摘要 · Abstract (English)
The capabilities of artificial intelligence systems have been advancing to a great extent, but these systems still struggle with failure modes, vulnerabilities, and biases. In this paper, we study the current state of the field, and present promising insights and perspectives regarding concerns that challenge the trustworthiness of AI models. In particular, this paper investigates the issues regarding three thrusts: safety, privacy, and bias, which hurt models' trustworthiness. For safety, we discuss safety alignment in the context of large language models, preventing them from generating toxic or harmful content. For bias, we focus on spurious biases that can mislead a network. Lastly, for privacy, we cover membership inference attacks in deep neural networks. The discussions addressed in this paper reflect our own experiments and observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。