提出可实时部署的安全与公平防护机制,提升智能系统的责任可信度。
Towards Responsible AI: Advances in Safety, Fairness, and Accountability of Autonomous Systems
- 设计抗延迟观测的安全屏障,适配真实驾驶环境
- 引入公平性防护,低干预成本下保障群体公平
- 构建意图量化框架,支持事故责任追溯
随着自主系统在关键社会领域的影响日益加深,确保人工智能的负责任使用变得至关重要。本文推进了人工智能在安全性、公平性、透明性和问责性方面的研究。在安全性方面,将经典确定性屏蔽技术扩展至对延迟观测具有鲁棒性的新形式,使其适用于实际部署;同时在模拟自动驾驶车辆中实现确定性和概率性安全屏障,有效防止与道路使用者发生碰撞。提出公平性屏蔽(fairness shields)这一新型后处理方法,用于在有限且周期性的时间范围内强化序列决策中的群体公平性,在严格满足公平约束的同时最小化干预成本。针对透明性与问责性,提出一个形式化框架以评估概率决策代理的意图行为,引入代理度与意图商等量化指标,可用于事后分析意图,帮助判定自主系统造成意外损害时的责任归属。最后,通过“反应式决策”框架统一前述成果,为可信AI研究奠定基础。整体贡献推动了更安全、更公平、更具问责性的智能系统发展。
原文摘要 · Abstract (English)
Ensuring responsible use of artificial intelligence (AI) has become imperative as autonomous systems increasingly influence critical societal domains. However, the concept of trustworthy AI remains broad and multi-faceted. This thesis advances knowledge in the safety, fairness, transparency, and accountability of AI systems. In safety, we extend classical deterministic shielding techniques to become resilient against delayed observations, enabling practical deployment in real-world conditions. We also implement both deterministic and probabilistic safety shields into simulated autonomous vehicles to prevent collisions with road users, validating the use of these techniques in realistic driving simulators. We introduce fairness shields, a novel post-processing approach to enforce group fairness in sequential decision-making settings over finite and periodic time horizons. By optimizing intervention costs while strictly ensuring fairness constraints, this method efficiently balances fairness with minimal interference. For transparency and accountability, we propose a formal framework for assessing intentional behaviour in probabilistic decision-making agents, introducing quantitative metrics of agency and intention quotient. We use these metrics to propose a retrospective analysis of intention, useful for determining responsibility when autonomous systems cause unintended harm. Finally, we unify these contributions through the ``reactive decision-making'' framework, providing a general formalization that consolidates previous approaches. Collectively, the advancements presented contribute practically to the realization of safer, fairer, and more accountable AI systems, laying the foundations for future research in trustworthy AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。