arXiv:2412.08515cs.LGcs.AI2024-12被引 3

让模型分类更透明,通过隐空间聚类提升可解释性

Enhancing Interpretability Through Loss-Defined Classification Objective in Structured Latent Spaces

  • 在隐空间引入距离度量学习,强制同类样本聚类
  • 分类准确率不变,但轮廓系数显著提高
  • 无需调参,适合各类数据集的可解释建模

监督学习通常基于数据驱动范式,模型参数自主优化以逼近真实标签,但缺乏显式规则或先验假设。尽管在多个基准数据集上取得成功,这类方法仍被视为黑箱,限制了可解释性与决策过程分析。本文提出 Latent Boost,将先进的距离度量学习融入监督分类任务,在保持分类性能的同时,使每类样本在中间层隐空间中的表示区域高度紧凑。利用隐空间的结构信息,该方法提升了分类可解释性(表现为更高的轮廓系数),并加速训练收敛。其性能优势仅需极少额外计算开销,适用于多种数据集且无需针对特定数据调整。该工作开创了一种新范式:将分类性能与模型透明性对齐,以应对黑箱模型挑战。

原文摘要 · Abstract (English)

Supervised machine learning often operates on the data-driven paradigm, wherein internal model parameters are autonomously optimized to converge predicted outputs with the ground truth, devoid of explicitly programming rules or a priori assumptions. Although data-driven methods have yielded notable successes across various benchmark datasets, they inherently treat models as opaque entities, thereby limiting their interpretability and yielding a lack of explanatory insights into their decision-making processes. In this work, we introduce Latent Boost, a novel approach that integrates advanced distance metric learning into supervised classification tasks, enhancing both interpretability and training efficiency. Thus during training, the model is not only optimized for classification metrics of the discrete data points but also adheres to the rule that the collective representation zones of each class should be sharply clustered. By leveraging the rich structural insights of intermediate model layer latent representations, Latent Boost improves classification interpretability, as demonstrated by higher Silhouette scores, while accelerating training convergence. These performance and latent structural benefits are achieved with minimum additional cost, making it broadly applicable across various datasets without requiring data-specific adjustments. Furthermore, Latent Boost introduces a new paradigm for aligning classification performance with improved model transparency to address the challenges of black-box models.

可解释性隐空间分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。