arXiv:2410.23880cs.LG2024-10

直接优化解释属性,实现可控的解释权衡。

Transparent Trade-offs between Properties of Explanations

  • 不依赖间接鼓励,直接优化解释的特定属性
  • 在属性冲突时能稳定生成符合需求的解释
  • 适合需要精准控制解释特性的应用场景

解释黑箱机器学习模型时,解释应具备某些理想属性。现有方法通过构造过程'鼓励'这些属性,但无法保证解释真正具备目标属性,也无法在属性冲突时进行优先级控制。本文提出直接优化解释以实现期望属性的方法,该方法不仅更一致地生成最优属性的解释,还允许用户调控不同属性间的权衡,从而根据具体任务定制所需解释。

原文摘要 · Abstract (English)

When explaining black-box machine learning models, it's often important for explanations to have certain desirable properties. Most existing methods `encourage' desirable properties in their construction of explanations. In this work, we demonstrate that these forms of encouragement do not consistently create explanations with the properties that are supposedly being targeted. Moreover, they do not allow for any control over which properties are prioritized when different properties are at odds with each other. We propose to directly optimize explanations for desired properties. Our direct approach not only produces explanations with optimal properties more consistently but also empowers users to control trade-offs between different properties, allowing them to create explanations with exactly what is needed for a particular task.

可解释性属性优化模型解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。