arXiv:2606.10669cs.LGcs.AI2026-06中稿 · as a position pape…

概念模型中的信息泄漏未必有害,合理利用反而能提升准确性与可干预性。

In Defense of Information Leakage in Concept-based Models

论文配图:In Defense of Information Leakage in Concept-based Models
图 1 · 摘自论文原文
  • 重新定义训练目标,主动引导良性泄漏以增强模型可解释性
  • 在概念不完整的真实场景中,适度泄漏有助于保持模型准确率
  • 适用于需要可干预、高精度的现实应用,如医疗诊断或自动驾驶

概念模型(CMs)通过与人类可理解的概念(如“圆形”、“条纹”)对齐表示来预测结果,但其表征常会泄露与概念无关的信息。传统观点认为此类泄漏有害,应彻底消除,以免影响模型可解释性。本文指出,这一观点缺乏充分证据支持,且在真实世界中会导致不切实际的模型设计。尤其在概念不完整的常见情况下,适度的信息泄漏反而是构建准确且可干预模型所必需的。为此,我们提出‘良性泄漏’的概念,并通过重构典型训练目标,使模型在不牺牲准确性和可干预性的前提下,主动鼓励并利用这种泄漏。

原文摘要 · Abstract (English)

Concept-based models (CMs), deep neural networks that ground their predictions on representations aligned with human-understandable concepts (e.g., "round", "stripes", etc.), have been shown to learn representations that leak concept-irrelevant information. As the traditional narrative goes, this leakage is undesirable and should be eradicated as it leads to uninterpretable models. In this paper, we posit that this conventional view of leakage in CMs is not only ill-posed, as the evidence of how leakage makes a model less interpretable is often inconclusive, but also bound to lead to impractical CMs under common real-world constraints. Specifically, we argue that in real-world settings where concept incompleteness is the norm, some leakage is often necessary for constructing accurate and intervenable CMs. To this end, we propose that there is such a thing as benign leakage and show that, by optimizing a reframing of the typical CM training objective, CMs can encourage and exploit this form of leakage without sacrificing accuracy or intervenability.

概念模型可解释性信息泄漏可干预性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。