构建可复现的算法救济评估框架,统一验证救济方法有效性
RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation

- 分五层模块化设计,支持灵活组合数据、模型与救济方法
- 对27种前沿救济方法进行可复现性分类,并验证其核心结论
- 提供交互式网页界面,支持多维度对比分析
算法救济方法为个体提供反事实解释,说明如何改变决策结果。尽管方法进展迅速,但系统性比较仍困难:现有框架难以扩展,缺乏互操作性和对方法宣称结果的严格验证。我们提出RecourseBench,一个围绕模块化、可复现性和交互性三大承诺构建的统一评估框架。该框架将流程分解为数据、预处理、模型、救济方法和评估五个完全解耦的层级,通过抽象接口与动态注册表管理。每个集成方法依据可用资源情况划分为四类可复现性等级,并对其核心实证主张进行系统验证。我们还提供交互式网页界面,支持基于配置的跨数据集、模型架构、方法和评估的灵活探索。据我们所知,RecourseBench是首个明确以结构化主张验证和数学严谨可复现性标准为基础的救济评估基准,同时包含最大规模的前沿救济算法集合(共27种)。
原文摘要 · Abstract (English)
Algorithmic recourse methods provide counterfactual explanations that inform individuals of the actions required to overturn an unfavorable model decision. Despite rapid methodological progress, principled comparison remains elusive; existing frameworks are often difficult to extend and lack both interoperability and systematic verification that integrated methods faithfully reproduce their originally reported claims. We introduce RecourseBench, a unified evaluation framework built around three commitments: modularity, reproducibility, and interactivity. The framework decomposes the pipeline into five fully decoupled layers---Data, Preprocessing, Model, Recourse Method, and Evaluation---governed by abstract interfaces and a dynamic registry. Every integrated method is classified into a four-tier reproducibility taxonomy based on artifact availability, followed by a systematic verification of its core empirical claims. We further provide an interactive web interface for flexible, configuration-driven exploration across datasets, model architectures, methods, and evaluations. To our knowledge, RecourseBench is the first benchmark to explicitly ground recourse evaluation in structured claim verification and mathematically rigorous reproducibility standards, all while featuring the largest collection of state-of-the-art recourse algorithms (27 in total).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。