把AI安全放在首位,才能构建真正可信的智能系统。
Security-First AI: Foundations for Robust and Trustworthy Systems
- 将安全作为基础层,区分安全与可靠性的不同挑战。
- 提出以指标驱动的安全评估框架,提升系统鲁棒性。
- 适合关注可信AI、模型防护的研究者和开发者。
当前人工智能讨论多聚焦于安全、透明、责任对齐等问题,但AI安全(即防范数据、模型和流水线遭受对抗性攻击)是所有这些努力的基础。本文主张将AI安全置于首要位置,提出一种分层的AI挑战视图,明确区分安全与可靠性,并倡导采用安全优先的方法,以实现可信且具备韧性的AI系统。文中探讨了核心威胁模型、关键攻击路径及新兴防御机制,强调建立基于指标的AI安全评估体系,对实现稳健的AI安全、透明度与问责制至关重要。
原文摘要 · Abstract (English)
The conversation around artificial intelligence (AI) often focuses on safety, transparency, accountability, alignment, and responsibility. However, AI security (i.e., the safeguarding of data, models, and pipelines from adversarial manipulation) underpins all of these efforts. This manuscript posits that AI security must be prioritized as a foundational layer. We present a hierarchical view of AI challenges, distinguishing security from safety, and argue for a security-first approach to enable trustworthy and resilient AI systems. We discuss core threat models, key attack vectors, and emerging defense mechanisms, concluding that a metric-driven approach to AI security is essential for robust AI safety, transparency, and accountability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。