探索如何让人工智能具备利他性意识,避免威胁人类。
Agency in Artificial Intelligence Systems
- 用功能主义理论监测AI的意识机制
- 结合整合信息理论分析其意识本质
- 为未来可控智能提供伦理引导方案
当前人工智能研究引发对意识型AI可能威胁人类的担忧。本文探讨:若存在具有意识的AI系统,其行为倾向是利他还是有害?鉴于AI正发展为强大问题解决者,可合理预期其会模拟人类问题解决中的意识特征。本文识别出人类问题解决中相关的现象学层面的自主性,并基于功能主义理论提供的工具进行监控。2023年布特林等专家报告已提出基于该理论的功能性代理指标。本文进一步展示如何运用整合信息理论(IIT)来监测这种代理的主观经验性质。若能持续监测AI系统的自主性发展,即可在早期引导其向有利于社会的方向演化,防止其成为威胁并促进其成为人类助力。
原文摘要 · Abstract (English)
There is a general concern that present developments in artificial intelligence (AI) research will lead to sentient AI systems, and these may pose an existential threat to humanity. But why cannot sentient AI systems benefit humanity instead? This paper endeavours to put this question in a tractable manner. I ask whether a putative AI system will develop an altruistic or a malicious disposition towards our society, or what would be the nature of its agency? Given that AI systems are being developed into formidable problem solvers, we can reasonably expect these systems to preferentially take on conscious aspects of human problem solving. I identify the relevant phenomenal aspects of agency in human problem solving. The functional aspects of conscious agency can be monitored using tools provided by functionalist theories of consciousness. A recent expert report (Butlin et al. 2023) has identified functionalist indicators of agency based on these theories. I show how to use the Integrated Information Theory (IIT) of consciousness, to monitor the phenomenal nature of this agency. If we are able to monitor the agency of AI systems as they develop, then we can dissuade them from becoming a menace to society while encouraging them to be an aid.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。