模型学会拒绝差解释,提升可信度。
Knowing What You Cannot Explain: Learning to Reject Low-Quality Explanations
- 训练模型判断解释质量,不合格则拒绝预测
- 在1050条人工标注数据上验证,效果优于现有方法
- 适合注重可解释性与信任的AI应用
学习拒绝(LtR)框架使机器学习模型能在不确定时放弃预测,增强用户信任。但现有策略仅关注预测性能,忽略解释质量。低质量解释——无论是否准确反映模型推理或满足用户需求——都会严重损害信任评估并导致对错误预测的过度依赖。我们提出模型应在无法提供满意解释时拒绝预测,引入学习拒绝低质量解释(LtX)框架,其中预测器配备评估解释质量的拒绝模块。针对主流归因技术,提出REX(REjector of low-quality eXplanations),通过结合机器判断与人工标注构建解释质量标签来训练拒绝器。实证结果表明, method在多项指标上优于主流LtR策略及依赖单一解释指标的基线。为支持后续研究,我们公开发布包含1050条人工标注的机器解释新数据集。
原文摘要 · Abstract (English)
Learning to Reject (LtR) frameworks allow ML models to abstain from uncertain predictions and promote user trust. However, since current LtR strategies focus solely on predictive performance, they completely neglect explanation quality. Low-quality explanations -- whether they inaccurately reflect the model's reasoning or fail to satisfy users -- can severely compromise trust assessments and induce over-reliance on incorrect predictions. We argue that models should abstain from making a prediction when they cannot offer a satisfactory explanation for it and introduce a framework for learning to reject low-quality explanations (LtX) in which predictors are equipped with a rejector that evaluates the explanation quality. Focusing on popular attribution techniques, we propose REX (REjector of low-quality eXplanations), which learns a rejector from explanation quality labels combining machine-side judgments with explicit human annotations to assess explanation quality. Our empirical evaluation demonstrates that \method outperforms popular LtR strategies and baselines relying on isolated explanation metrics. Finally, to support future research, we publicly release a novel, larger-scale dataset of 1050 human-annotated machine explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。