从行为推断智能体信念有理论极限,影响安全与公平评估
The Limits of Predicting Agents from Behaviour
- 基于世界模型假设,推导行为预测的理论上限
- 在未见环境中的行为预测存在可计算的可靠边界
- 适用于安全、公平等需解释性推理的AI研究
随着AI系统复杂性及其与世界互动的增加,对其行为提供解释对安全部署至关重要。对于智能体,最自然的预测抽象是赋予其信念、意图和目标。若智能体表现出具有特定目标或信念的行为,则可在新情境下做出合理预测,包括那些难以进行全面安全评估的情形。本文在假设智能体行为由世界模型引导的前提下,精确回答了如何从行为中推断信念,以及这些推断信念在新环境中的预测可靠性问题。我们的贡献是推导出智能体在新(未见)部署环境中的行为理论界限,这代表了仅从行为数据预测意图型智能体的理论极限。我们讨论了这些结果对公平性和安全性等多个研究领域的启示。
原文摘要 · Abstract (English)
As the complexity of AI systems and their interactions with the world increases, generating explanations for their behaviour is important for safely deploying AI. For agents, the most natural abstractions for predicting behaviour attribute beliefs, intentions and goals to the system. If an agent behaves as if it has a certain goal or belief, then we can make reasonable predictions about how it will behave in novel situations, including those where comprehensive safety evaluations are untenable. How well can we infer an agent's beliefs from their behaviour, and how reliably can these inferred beliefs predict the agent's behaviour in novel situations? We provide a precise answer to this question under the assumption that the agent's behaviour is guided by a world model. Our contribution is the derivation of novel bounds on the agent's behaviour in new (unseen) deployment environments, which represent a theoretical limit for predicting intentional agents from behavioural data alone. We discuss the implications of these results for several research areas including fairness and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。