剖析概念模型的潜在风险,揭示其在医疗金融等关键场景中的可信隐患。
A Comprehensive Survey on the Risks and Limitations of Concept-based Models
- 系统梳理概念模型在结构与训练中的固有缺陷
- 指出概念泄露、表征纠缠等问题影响模型可靠性
- 适合关注可解释AI安全与落地挑战的研究者
概念模型是一类具备内在可解释性的神经网络,通过人类可理解的'概念'解释预测结果,已在医疗诊断和金融风控等关键领域取得成功。然而,近期研究揭示其在结构设计、训练过程、基础假设及对抗脆弱性方面存在显著局限。例如,概念泄露、表征纠缠以及对扰动的鲁棒性不足,均影响其可靠性和泛化能力。此外,人工干预的有效性仍存疑,制约其真实应用场景。本文全面综述概念模型面临的风险与限制,聚焦监督与无监督范式中常见挑战及其缓解架构,并探讨提升可靠性的最新进展,提出开放问题与未来研究方向。
原文摘要 · Abstract (English)
Concept-based Models are a class of inherently explainable networks that improve upon standard Deep Neural Networks by providing a rationale behind their predictions using human-understandable `concepts'. With these models being highly successful in critical applications like medical diagnosis and financial risk prediction, there is a natural push toward their wider adoption in sensitive domains to instill greater trust among diverse stakeholders. However, recent research has uncovered significant limitations in the structure of such networks, their training procedure, underlying assumptions, and their susceptibility to adversarial vulnerabilities. In particular, issues such as concept leakage, entangled representations, and limited robustness to perturbations pose challenges to their reliability and generalization. Additionally, the effectiveness of human interventions in these models remains an open question, raising concerns about their real-world applicability. In this paper, we provide a comprehensive survey on the risks and limitations associated with Concept-based Models. In particular, we focus on aggregating commonly encountered challenges and the architecture choices mitigating these challenges for Supervised and Unsupervised paradigms. We also examine recent advances in improving their reliability and discuss open problems and promising avenues of future research in this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。