承认伦理指标矛盾,反而能构建更负责任的AI系统
Embracing Contradiction: Theoretical Inconsistency Will Not Impede the Road of Building Responsible AI Systems
- 将矛盾指标视为多元价值体现,而非需消除的缺陷
- 多指标并行提升对复杂伦理概念的信息保真度
- 联合优化矛盾目标可增强模型泛化与鲁棒性
本文主张,负责任AI(RAI)指标间常见的理论不一致(如公平性定义差异或准确率与隐私间的权衡)不应被视作缺陷,而应被接纳为宝贵特性。通过将指标视为相互冲突的目标,可带来三重益处:(1) 规范多元性:保留所有潜在矛盾指标,确保RAI中蕴含的多样道德立场与利益相关者价值得到充分代表;(2) 认知完整性:使用多组有时冲突的指标,能更全面捕捉复杂的伦理概念,从而比单一简化定义保留更多概念信息;(3) 隐式正则化:联合优化理论冲突目标可防止模型过度拟合单一指标,引导模型走向更具泛化能力与抗现实复杂性的解。相反,强制统一指标的做法可能压缩价值多样性、丧失概念深度并降低模型性能。因此我们呼吁转变RAI理论与实践:从回避不一致转向界定可接受的不一致阈值,并阐明实践中实现近似一致的机制。
原文摘要 · Abstract (English)
This position paper argues that the theoretical inconsistency often observed among Responsible AI (RAI) metrics, such as differing fairness definitions or tradeoffs between accuracy and privacy, should be embraced as a valuable feature rather than a flaw to be eliminated. We contend that navigating these inconsistencies, by treating metrics as divergent objectives, yields three key benefits: (1) Normative Pluralism: Maintaining a full suite of potentially contradictory metrics ensures that the diverse moral stances and stakeholder values inherent in RAI are adequately represented. (2) Epistemological Completeness: The use of multiple, sometimes conflicting, metrics allows for a more comprehensive capture of multifaceted ethical concepts, thereby preserving greater informational fidelity about these concepts than any single, simplified definition. (3) Implicit Regularization: Jointly optimizing for theoretically conflicting objectives discourages overfitting to one specific metric, steering models towards solutions with enhanced generalization and robustness under real-world complexities. In contrast, efforts to enforce theoretical consistency by simplifying or pruning metrics risk narrowing this value diversity, losing conceptual depth, and degrading model performance. We therefore advocate for a shift in RAI theory and practice: from getting trapped in inconsistency to characterizing acceptable inconsistency thresholds and elucidating the mechanisms that permit robust, approximated consistency in practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。