揭示AI安全与对齐的理论极限,警示其内在不可靠性。
Robust AI Security and Alignment: A Sisyphean Endeavor?
- 用哥德尔不完备性定理拓展出AI安全的理论瓶颈
- 证明了任何AI系统都无法完全保证安全与对齐
- 为追求可靠AI的开发者提供关键预警和应对思路
本文通过将哥德尔不完备性定理扩展至人工智能领域,建立了AI安全与对齐鲁棒性的信息论限制。认识到这些限制并提前应对由此带来的挑战,对于负责任地采用AI技术至关重要。文章还提出了应对这些挑战的实际方法,并进一步证明了AI系统在认知推理方面的根本局限性。
原文摘要 · Abstract (English)
This manuscript establishes information-theoretic limitations for robustness of AI security and alignment by extending Gödel's incompleteness theorem to AI. Knowing these limitations and preparing for the challenges they bring is critically important for the responsible adoption of the AI technology. Practical approaches to dealing with these challenges are provided as well. Broader implications for cognitive reasoning limitations of AI systems are also proven.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。