arXiv:2504.07170cs.LGcs.AI2025-04被引 5

提出可信AI需统筹考虑各属性间的相互影响,避免片面优化

Trustworthy AI Must Account for Interactions

  • 系统分析公平性、隐私等五类可信属性间的负面互斥关系
  • 实证指出差分隐私可能加剧模型偏见,破坏公平性
  • 倡导整体性研究范式,适合政策制定与跨领域应用者

可信AI旨在协同实现公平性、隐私保护、鲁棒性、可解释性和不确定性量化等多重目标。然而,提升某一属性常引发其他属性的退化。本文梳理五类属性的典型方法,系统考察其两两组合中的负向交互:例如差分隐私虽增强隐私,却可能放大模型偏见,损害公平性。基于此,我们主张当前孤立优化单个或少数属性的研究范式不足,可信AI研究必须同时考量所有相关维度的交互影响。为此,我们提供整合信任的实践指引,举出金融行业中的交互案例,并讨论替代视角。

原文摘要 · Abstract (English)

Trustworthy AI encompasses many aspirational aspects for aligning AI systems with human values, including fairness, privacy, robustness, explainability, and uncertainty quantification. Ultimately the goal of Trustworthy AI research is to achieve all aspects simultaneously. However, efforts to enhance one aspect often introduce unintended trade-offs that negatively impact others. In this position paper, we review notable approaches to these five aspects and systematically consider every pair, detailing the negative interactions that can arise. For example, applying differential privacy to model training can amplify biases, undermining fairness. Drawing on these findings, we take the position that current research practices of improving one or two aspects in isolation are insufficient. Instead, research on Trustworthy AI must account for interactions between aspects and adopt a holistic view across all relevant axes at once. To illustrate our perspective, we provide guidance on how practitioners can work towards integrated trust, examples of how interactions affect the financial industry, and alternative views.

可信AI属性交互公平性隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。