arXiv:2409.13684cs.LGcs.AI2024-09被引 5

提出可衡量特征与专家知识对齐程度的基准,发现现有解释方法效果不佳。

The FIX Benchmark: Extracting Features Interpretable to eXperts

  • 构建跨领域专家知识对齐度评估框架,支持多模态数据
  • 实测主流解释方法与专家指定特征对齐度低
  • 适用于天体物理、心理学、医学等需专家理解的场景

基于特征的方法常用于解释模型预测,但通常隐含假设可解释特征已存在。然而在高维数据中,甚至领域专家也难以数学定义重要特征。能否自动提取与专家知识一致的特征集合?为此,我们提出FIX(Features Interpretable to eXperts)基准,用于衡量特征集合与专家知识的对齐程度。通过与领域专家合作,我们提出了适用于天体物理、心理学、医学等多个现实场景的统一评估指标FIXScore,涵盖视觉、语言和时间序列数据模态。利用FIXScore,我们发现主流特征解释方法与专家指定的知识对齐度较差,凸显了开发更优专家可解释特征识别方法的必要性。

原文摘要 · Abstract (English)

Feature-based methods are commonly used to explain model predictions, but these methods often implicitly assume that interpretable features are readily available. However, this is often not the case for high-dimensional data, and it can be hard even for domain experts to mathematically specify which features are important. Can we instead automatically extract collections or groups of features that are aligned with expert knowledge? To address this gap, we present FIX (Features Interpretable to eXperts), a benchmark for measuring how well a collection of features aligns with expert knowledge. In collaboration with domain experts, we propose FIXScore, a unified expert alignment measure applicable to diverse real-world settings across cosmology, psychology, and medicine domains in vision, language, and time series data modalities. With FIXScore, we find that popular feature-based explanation methods have poor alignment with expert-specified knowledge, highlighting the need for new methods that can better identify features interpretable to experts.

可解释性专家对齐特征提取基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。