用因果关系破解可信AI的多目标冲突难题
Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

- 将可信AI的矛盾目标重新解释为数据生成过程中的不兼容不变性要求
- 通过案例和模拟证明因果框架能化解性能与公平、鲁棒等目标的权衡
- 为传统机器学习和大模型提供统一的可信性优化思路,适合研究者参考
随着人工智能(包括机器学习模型和基础模型)在高风险领域日益普及,确保其可信性已成为核心挑战。然而,公平性、鲁棒性、隐私性和可解释性等核心可信目标难以同时实现,尤其在保持实用性的情况下。本文认为,因果关系是理解并平衡可信AI多目标间权衡的关键。我们通过重新诠释这些权衡为数据生成过程变化下的不兼容不变性要求来论证这一观点。结合文献中的案例分析和一个简化的合成数据模拟,我们展示了因果框架如何统一理解可信AI中权衡的成因,并通过选择性不变性来缓解或解决这些矛盾。该视角适用于经典机器学习模型和大规模基础模型。最后,我们提出了利用因果方法构建既可信又高性能的AI所面临的开放挑战与机遇。
原文摘要 · Abstract (English)
As artificial intelligence (AI), including machine learning (ML) models and foundation models (FMs), are increasingly deployed in high-stakes domains, ensuring their trustworthiness has become a central challenge. However, the core trustworthy AI objectives, such as fairness, robustness, privacy, and explainability, are hard to achieve simultaneously, especially while preserving utility. This position paper argues that causality is necessary to understand and balance trade-offs in performance and multiple objectives of trustworthy AI. We ground our arguments in re-interpreting trustworthy AI trade-offs as incompatible invariance requirements under different changes to the data-generating process. We then illustrate this argument through case-study analyses from the literature and a stylized synthetic-data simulation, showing that causality provides a unifying framework for understanding how trade-offs in trustworthy AI arise and how they can be softened or resolved through selective invariance. This perspective applies to both classical ML models and large-scale FMs. Finally, we outline open challenges and opportunities for using causality to build both trustworthy and high-performing AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。