Quanda工具包统一评估训练数据归属方法,助力模型可解释性研究
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
- 提供统一接口,兼容多种TDA方法实现
- 内置多维度评估指标,支持系统化对比分析
- 开源易用,适合研究人员快速验证方法性能
近年来,训练数据归属(TDA)方法成为神经网络可解释性的新兴方向。尽管相关研究蓬勃发展,但对归属结果的评估仍缺乏系统性。类似传统特征归因评估指标的发展,已有若干独立指标被提出用于在不同场景下评估TDA方法的质量。然而,由于缺乏统一框架进行系统比较,导致TDA方法的可信度受限,难以广泛应用。为此,我们提出Quanda——一个Python工具包,旨在促进TDA方法的评估。除提供全面的评估指标外,Quanda还为不同仓库中的现有TDA实现提供统一接口,支持无缝集成与系统基准测试。该工具包用户友好、经过充分测试、文档完善,已作为开源库发布于PyPI及https://github.com/dilyabareeva/quanda。
原文摘要 · Abstract (English)
In recent years, training data attribution (TDA) methods have emerged as a promising direction for the interpretability of neural networks. While research around TDA is thriving, limited effort has been dedicated to the evaluation of attributions. Similar to the development of evaluation metrics for traditional feature attribution approaches, several standalone metrics have been proposed to evaluate the quality of TDA methods across various contexts. However, the lack of a unified framework that allows for systematic comparison limits trust in TDA methods and stunts their widespread adoption. To address this research gap, we introduce Quanda, a Python toolkit designed to facilitate the evaluation of TDA methods. Beyond offering a comprehensive set of evaluation metrics, Quanda provides a uniform interface for seamless integration with existing TDA implementations across different repositories, thus enabling systematic benchmarking. The toolkit is user-friendly, thoroughly tested, well-documented, and available as an open-source library on PyPi and under https://github.com/dilyabareeva/quanda.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。