提出可信赖AI框架,兼顾模型可解释性与公平、隐私、因果等伦理属性。
Towards the Formalization of a Trustworthy AI for Mining Interpretable Models explOiting Sophisticated Algorithms
- 构建MIMOSA框架,统一多类型数据的可解释建模方法。
- 定义公平、隐私、因果三类伦理属性的评估指标与验证流程。
- 适用于医疗、金融等需高可信度决策的场景。
可解释性设计模型对提升自动化决策系统的信任度、责任性与安全性至关重要。本文正式化了MIMOSA(Mining Interpretable Models explOiting Sophisticated Algorithms)框架的基础,该框架是一种综合性方法论,旨在生成在可解释性与性能之间取得平衡,并嵌入关键伦理属性的预测模型。我们形式化了涵盖表格数据、时间序列、图像、文本、交易记录和轨迹等多种任务与数据类型的监督学习设定。分析了三类主要可解释模型:特征重要性、规则与实例基模型,分别探讨其可解释维度、推理机制与复杂度。除可解释性外,本文还形式化定义了因果性、公平性与隐私三大关键伦理属性,提供相应的形式化定义、评估指标与验证程序。进一步分析这些属性间的内在权衡,探讨如何在可解释建模流程中嵌入隐私要求、公平约束与因果推理。通过在模型生成过程中评估伦理指标,该框架为构建不仅准确、可解释,且公平、隐私保护、具备因果意识的可信AI系统奠定了理论基础。
原文摘要 · Abstract (English)
Interpretable-by-design models are crucial for fostering trust, accountability, and safe adoption of automated decision-making models in real-world applications. In this paper we formalize the ground for the MIMOSA (Mining Interpretable Models explOiting Sophisticated Algorithms) framework, a comprehensive methodology for generating predictive models that balance interpretability with performance while embedding key ethical properties. We formally define here the supervised learning setting across diverse decision-making tasks and data types, including tabular data, time series, images, text, transactions, and trajectories. We characterize three major families of interpretable models: feature importance, rule, and instance based models. For each family, we analyze their interpretability dimensions, reasoning mechanisms, and complexity. Beyond interpretability, we formalize three critical ethical properties, namely causality, fairness, and privacy, providing formal definitions, evaluation metrics, and verification procedures for each. We then examine the inherent trade-offs between these properties and discuss how privacy requirements, fairness constraints, and causal reasoning can be embedded within interpretable pipelines. By evaluating ethical measures during model generation, this framework establishes the theoretical foundations for developing AI systems that are not only accurate and interpretable but also fair, privacy-preserving, and causally aware, i.e., trustworthy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。