arXiv:2411.01956cs.LGcs.CY2024-11

用用户偏好筛选解释模型,让机器解释更可信且一致。

EXAGREE: Mitigating Explanation Disagreement with Stakeholder-Aligned Models

  • 从多个性能相近模型中选最符合人类偏好的解释模型。
  • 在6个真实数据集上同时提升解释忠实性、合理性和公平性。
  • 适合需要可解释性与信任的医疗、金融等高风险领域。

由于不同归因方法或模型内部机制导致的解释冲突,限制了机器学习在安全关键领域的应用。本文将这种分歧转化为优势,提出EXAGREE框架:通过两阶段方法从一组表现相似的模型中选择一个利益相关方对齐解释模型(SAEM),以最大化利益相关方-机器一致性(SMA)——该指标统一了忠实性与合理性。EXAGREE结合可微分掩码归因网络(DMAN)与单调可微排序,实现对受限模型空间内的梯度搜索。在六个真实世界数据集上的实验表明,该方法在保持任务准确性的前提下,相比基线模型在忠实性、合理性和公平性上均有显著提升。大量消融实验、显著性检验与案例研究验证了方法在实际应用中的鲁棒性与可行性。

原文摘要 · Abstract (English)

Conflicting explanations, arising from different attribution methods or model internals, limit the adoption of machine learning models in safety-critical domains. We turn this disagreement into an advantage and introduce EXplanation AGREEment (EXAGREE), a two-stage framework that selects a Stakeholder-Aligned Explanation Model (SAEM) from a set of similar-performing models. The selection maximizes Stakeholder-Machine Agreement (SMA), a single metric that unifies faithfulness and plausibility. EXAGREE couples a differentiable mask-based attribution network (DMAN) with monotone differentiable sorting, enabling gradient-based search inside the constrained model space. Experiments on six real-world datasets demonstrate simultaneous gains of faithfulness, plausibility, and fairness over baselines, while preserving task accuracy. Extensive ablation studies, significance tests, and case studies confirm the robustness and feasibility of the method in practice.

可解释性模型选择可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。