arXiv:2410.06946cs.AI2024-10

提出AI安全三框架,用三个可落地方向提升可靠性。

A Trilogy of AI Safety Frameworks: Paths from Facts and Knowledge Gaps to Reliable Predictions and New Knowledge

  • 构建三阶段安全框架,聚焦可实现的改进路径
  • 在生物医学机器学习中验证了概念可行性
  • 适合关注安全与创新平衡的研究者和从业者

人工智能安全已成为科学界内外广泛关注的前沿议题。从存在性风险到深度伪造和机器学习中的偏见,潜在威胁范围广泛。本文将庞大的AI安全挑战简化为三个重要且可实现的突破机会,这些方向有望在不抑制关键领域创新的前提下,短期内提升AI的安全性与可靠性。基于生物医学领域多个已验证的案例研究,本论文展示了该愿景的可行性。

原文摘要 · Abstract (English)

AI Safety has become a vital front-line concern of many scientists within and outside the AI community. There are many immediate and long term anticipated risks that range from existential risk to human existence to deep fakes and bias in machine learning systems [1-5]. In this paper, we reduce the full scope and immense complexity of AI safety concerns to a trilogy of three important but tractable opportunities for advances that have the short-term potential to improve AI safety and reliability without reducing AI innovation in critical domains. In this perspective, we discuss this vision based on several case studies that already produced proofs of concept in critical ML applications in biomedical science.

AI安全框架设计可信预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。