arXiv:2507.12192cs.LG2025-07被引 2

让证据聚类结果可解释,用决策树给出可信的错误容忍解释。

Explainable Evidential Clustering

  • 基于证据理论设计可解释的聚类方法,引入容错机制。
  • 提出迭代误判最小化算法,解释满意率达93%。
  • 适合医疗等高风险领域,支持决策偏好定制。

无监督分类是机器学习的基础问题。现实数据常含不确定性与模糊性,传统方法难以应对。基于达姆斯特定理的证据聚类能更好处理此类问题。本文研究证据聚类结果的可解释性,这对医疗等高风险领域至关重要。分析表明,在一般情况下,代表性是决策树作为溯因解释器的充要条件。基于此,我们推广该概念以支持部分标注,通过效用函数表示‘可接受的错误’,定义证据错误为解释成本,并构建适配证据分类器的解释器。最后,提出迭代证据误判最小化(IEMM)算法,为证据聚类提供可解释且谨慎的决策树解释。在合成与真实数据上验证了该算法,结合决策者偏好后,解释满意度达93%。

原文摘要 · Abstract (English)

Unsupervised classification is a fundamental machine learning problem. Real-world data often contain imperfections, characterized by uncertainty and imprecision, which are not well handled by traditional methods. Evidential clustering, based on Dempster-Shafer theory, addresses these challenges. This paper explores the underexplored problem of explaining evidential clustering results, which is crucial for high-stakes domains such as healthcare. Our analysis shows that, in the general case, representativity is a necessary and sufficient condition for decision trees to serve as abductive explainers. Building on the concept of representativity, we generalize this idea to accommodate partial labeling through utility functions. These functions enable the representation of "tolerable" mistakes, leading to the definition of evidential mistakeness as explanation cost and the construction of explainers tailored to evidential classifiers. Finally, we propose the Iterative Evidential Mistake Minimization (IEMM) algorithm, which provides interpretable and cautious decision tree explanations for evidential clustering functions. We validate the proposed algorithm on synthetic and real-world data. Taking into account the decision-maker's preferences, we were able to provide an explanation that was satisfactory up to 93% of the time.

聚类解释证据理论可解释性决策树

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。