让算法推荐更可靠:用户可指定信心目标,生成更稳健的反事实建议。
Target-confidence Recourse Using tSeTlin machines: TRUST

- 直接搜索满足用户设定信心阈值的最小改动,避免边界附近脆弱解。
- 在多个数据集上实现完美鲁棒性,如哈伯曼数据集0.92信心下L2距离仅0.10。
- 通过规则稳定性揭示决策是牢固还是脆弱,适合高风险场景使用。
反事实解释广泛用于高风险决策系统中的算法救济。现有方法通常寻找使模型预测翻转的最小输入变化,但决策者不仅关注预测标签,还依赖置信度阈值和风险边际。仅勉强越过决策边界的反事实在噪声或模型波动下可能不稳定。本文提出目标置信度救济框架TRUST,允许用户显式指定期望的预测置信度。不同于先生成反事实再评估置信度,TRUST直接搜索满足用户定义置信目标的最小变化,可对比不同救济方案在成本、置信度与鲁棒性上的表现。我们结合概率型Tsetlin机(PTM)与贝叶斯优化实例化TRUST。PTM基于概率的条款结构将预测置信度与决策规则稳定性关联。实验表明,满足相同规则的反事实仍可能因条款激活的稳固程度不同而可靠性差异显著。在合成与真实数据集上,目标置信反事实相比传统边界方法更具鲁棒性与可解释性。在多个基准测试中,TRUST实现完美鲁棒性,同时保持低救济成本,例如在哈伯曼数据集上0.92置信度下L2距离仅为0.10。通过显式控制置信度并暴露规则级稳定性,TRUST为高风险决策支持提供可操作的救济方案。
原文摘要 · Abstract (English)
Counterfactual explanations are widely used to provide algorithmic recourse in high-stakes decision-making systems. Most existing methods seek the smallest change to an input that flips a model's decision. However, decision-makers often rely not only on predicted labels but also on confidence thresholds and risk margins. Counterfactuals that barely cross a decision boundary can be fragile and unstable under noise or model variation. In this paper, we propose Target-confidence Recourse Using tSeTlin machines (TRUST), a framework in which users explicitly specify the desired prediction confidence for recourse. Rather than generating counterfactuals and evaluating confidence afterward, TRUST directly searches for minimal changes that satisfy a user-defined confidence target, enabling comparison of recourse options in terms of cost, confidence, and robustness. We instantiate TRUST using a Probabilistic Tsetlin Machine (PTM) combined with Bayesian optimization. The probabilistic clause-based structure of PTM links prediction confidence to the stability of decision rules. We show that counterfactuals satisfying the same rules can still differ substantially in reliability depending on how securely they satisfy those rules, revealing whether decisions are supported by robust or fragile clause activations. Experiments on synthetic and real-world datasets demonstrate that target-confidence counterfactuals produce more robust and interpretable recourse than conventional boundary-based approaches. Across multiple benchmarks, TRUST achieves perfect robustness while maintaining low recourse cost, including an L2 distance of 0.10 on the Haberman dataset at 0.92 confidence. By explicitly controlling confidence and exposing rule-level stability, TRUST provides actionable recourse for high-stakes decision support.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。