arXiv:2605.28626cs.LG2026-05中稿 · AAAI

混合可解释模型可能对不同群体分配不均的解释权,需警惕算法偏见。

When Interpretability Is Unequally Distributed: Fairness in Hybrid Interpretable Models

  • 设计混合模型时按人群分配解释性决策,引发公平性问题。
  • 实验发现中等透明度下群体间解释覆盖差异显著,最高达27%。
  • 加入简单约束即可大幅降低差异,且不影响准确率与模型简洁性。

混合可解释模型通过将部分样本分配给透明组件、其余交由黑箱模型处理,实现精度与可解释性的灵活权衡。然而,这种设计也带来新的程序公平性问题:某些人口群体可能系统性地获得可解释决策,而其他群体则被过度引导至黑箱。本文将此问题形式化为解释覆盖差异(ICD),一种基于群体平等性的路由决策衡量标准。利用预测多重性工具,在四种混合可解释学习方法、三个标准公平性基准数据集及多个敏感属性上研究ICD。实验表明,在中间透明度区间(即透明与黑箱组件均活跃时),存在显著的解释覆盖差异。进一步证明,简单的覆盖率差异约束可在精确混合学习方法中显著降低ICD,对准确率和稀疏性影响微小。在若干场景下,缓解ICD还提升了传统算法公平性指标。结果表明,混合可解释模型不仅需评估预测公平性,还应审查其在个体与群体间如何分配解释性。

原文摘要 · Abstract (English)

Hybrid interpretable models combine a transparent component with a black-box model by assigning some examples to the former and deferring the rest to the latter. While this design enables flexible tradeoffs between accuracy and interpretability, it also raises a distinct procedural fairness concern: some demographic groups may systematically receive interpretable decisions, while others are disproportionately routed to a black box. We formalize this issue as Interpretability Coverage Disparity (ICD), a demographic-parity-style measure applied to the routing decision of hybrid interpretable models. Using tools from predictive multiplicity, we study ICD across four hybrid interpretable learning methods, three standard fairness benchmark datasets, and multiple sensitive attributes. Our experiments reveal substantial ICD in intermediate transparency regimes, where both the interpretable and black-box components are actively used. We further show that simple coverage-disparity constraints can significantly reduce ICD in exact hybrid learning methods, with marginal impact on accuracy and sparsity. In several settings, ICD mitigation also improves standard algorithmic fairness metrics. These results show that hybrid interpretable models should be audited not only for predictive fairness, but also for how they allocate interpretability across individuals and groups.

可解释性公平性混合模型算法偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。