用概念学习自动解释知识图问答系统表现,提升评估可读性
Explainable Benchmarking through the Lense of Concept Learning
- 通过新方法PruneCEL从知识图谱中学习可解释概念
- 在可解释基准测试中F1最高提升0.55分,优于现有方法
- 80%用户能根据解释准确预测系统行为,适合评估优化者
科学方法中,系统对比评估(基准测试)至关重要,但性能常被少数指标概括,细节分析耗时且易受偏见影响。本文提出可解释基准测试新范式,旨在自动生成系统表现的解释。以基于知识图谱的问答系统为例,采用名为PruneCEL的新概念学习方法生成解释。实验表明,PruneCEL在可解释基准测试任务中,F1值相较当前最优方法最高提升0.55。一项包含41名参与者的任务驱动型用户研究显示,在80%情况下,多数参与者能准确依据解释预测系统行为。代码与数据已公开于https://github.com/dice-group/PruneCEL/tree/K-cap2025。
原文摘要 · Abstract (English)
Evaluating competing systems in a comparable way, i.e., benchmarking them, is an undeniable pillar of the scientific method. However, system performance is often summarized via a small number of metrics. The analysis of the evaluation details and the derivation of insights for further development or use remains a tedious manual task with often biased results. Thus, this paper argues for a new type of benchmarking, which is dubbed explainable benchmarking. The aim of explainable benchmarking approaches is to automatically generate explanations for the performance of systems in a benchmark. We provide a first instantiation of this paradigm for knowledge-graph-based question answering systems. We compute explanations by using a novel concept learning approach developed for large knowledge graphs called PruneCEL. Our evaluation shows that PruneCEL outperforms state-of-the-art concept learners on the task of explainable benchmarking by up to 0.55 points F1 measure. A task-driven user study with 41 participants shows that in 80\% of the cases, the majority of participants can accurately predict the behavior of a system based on our explanations. Our code and data are available at https://github.com/dice-group/PruneCEL/tree/K-cap2025
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。