arXiv:2505.02313cs.CYcs.AI2025-05被引 4

厘清AI安全的定义,主张所有防害研究都属安全范畴

What Is AI Safety? What Do We Want It to Be?

  • 以防止或减轻AI危害为核心定义AI安全
  • 涵盖偏见、虚假信息、隐私等长期边缘议题
  • 适合关注AI风险治理与跨领域安全的研究者

AI安全旨在预防或降低人工智能系统带来的损害。当前主流观点认为,该领域核心在于防范未来系统可能引发的灾难性风险,并将其视为安全工程的分支。本文通过概念工程方法指出,这种趋势并不理想。我们主张应采纳‘安全观’——即任何旨在防止或减少AI危害的研究都属于AI安全范畴。这一定义在描述上可统一传统核心与边缘议题(如偏见、虚假信息、隐私),在规范上则避免人为划分,仅依据措施有效性评估各类防害工作。

原文摘要 · Abstract (English)

The field of AI safety seeks to prevent or reduce the harms caused by AI systems. A simple and appealing account of what is distinctive of AI safety as a field holds that this feature is constitutive: a research project falls within the purview of AI safety just in case it aims to prevent or reduce the harms caused by AI systems. Call this appealingly simple account The Safety Conception of AI safety. Despite its simplicity and appeal, we argue that The Safety Conception is in tension with at least two trends in the ways AI safety researchers and organizations think and talk about AI safety: first, a tendency to characterize the goal of AI safety research in terms of catastrophic risks from future systems; second, the increasingly popular idea that AI safety can be thought of as a branch of safety engineering. Adopting the methodology of conceptual engineering, we argue that these trends are unfortunate: when we consider what concept of AI safety it would be best to have, there are compelling reasons to think that The Safety Conception is the answer. Descriptively, The Safety Conception allows us to see how work on topics that have historically been treated as central to the field of AI safety is continuous with work on topics that have historically been treated as more marginal, like bias, misinformation, and privacy. Normatively, taking The Safety Conception seriously means approaching all efforts to prevent or mitigate harms from AI systems based on their merits rather than drawing arbitrary distinctions between them.

AI安全概念定义风险治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。