对比多种方法,发现解释性AI能更好解决模型依赖虚假关联的问题。
Reproducibility study on how to find Spurious Correlations, Shortcut Learning, Clever Hans or Group-Distributional non-robustness and how to fix them
- 用可解释AI方法对比传统方法,统一不同领域对模型缺陷的应对思路。
- 基于因果反事实的知识蒸馏(CFKD)在提升泛化性能上最稳定有效。
- 多数方法受限于需标注分组标签,小样本和不平衡数据让模型调优失效。
深度神经网络在医疗诊断、自动驾驶等高风险领域应用日益广泛,但其可靠性保障研究因术语割裂而分散。尽管分布鲁棒优化(DRO)、不变风险最小化(IRM)、捷径学习、简单性偏见及聪明汉斯效应均旨在避免模型依赖虚假相关,但各领域常互不引用。本可复现研究通过在有限数据与严重子群体不平衡条件下,对比基于可解释人工智能(XAI)的方法与主流非XAI基线。实验表明,XAI方法整体优于非XAI,其中反事实知识蒸馏(CFKD)表现最一致。同时发现,多数方法依赖分组标签,人工标注难实现,自动工具如谱相关分析(SpRAy)在复杂特征与严重不平衡下失效;更严重的是,验证集中少数群体样本稀缺,导致模型选择与超参数调优不可靠,制约了安全关键场景中稳健模型的部署。
原文摘要 · Abstract (English)
Deep Neural Networks (DNNs) are increasingly utilized in high-stakes domains like medical diagnostics and autonomous driving where model reliability is critical. However, the research landscape for ensuring this reliability is terminologically fractured across communities that pursue the same goal of ensuring models rely on causally relevant features rather than confounding signals. While frameworks such as distributionally robust optimization (DRO), invariant risk minimization (IRM), shortcut learning, simplicity bias, and the Clever Hans effect all address model failure due to spurious correlations, researchers typically only reference work within their own domains. This reproducibility study unifies these perspectives through a comparative analysis of correction methods under challenging constraints like limited data availability and severe subgroup imbalance. We evaluate recently proposed correction methods based on explainable artificial intelligence (XAI) techniques alongside popular non-XAI baselines using both synthetic and real-world datasets. Findings show that XAI-based methods generally outperform non-XAI approaches, with Counterfactual Knowledge Distillation (CFKD) proving most consistently effective at improving generalization. Our experiments also reveal that the practical application of many methods is hindered by a dependency on group labels, as manual annotation is often infeasible and automated tools like Spectral Relevance Analysis (SpRAy) struggle with complex features and severe imbalance. Furthermore, the scarcity of minority group samples in validation sets renders model selection and hyperparameter tuning unreliable, posing a significant obstacle to the deployment of robust and trustworthy models in safety-critical areas.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。