揭秘自动聚类中数据特征对决策的影响,提升自动化系统可解释性。
Explaining AutoClustering: Uncovering Meta-Feature Contribution in AutoML for Clustering
- 构建元特征分类体系,系统分析22种自动聚类方法的特征使用模式。
- 发现关键元特征在不同算法选择中的稳定重要性,揭示现有策略的结构性缺陷。
- 结合全局与局部解释工具,为可解释的自动化机器学习设计提供实操指南。
自动聚类方法旨在通过数据集元特征上的元学习,自动化无监督学习任务,包括算法选择(AS)、超参数优化(HPO)和流程合成(PS)。尽管这些系统表现优异,其推荐结果常难以解释:数据集元特征对算法与超参数选择的影响通常不透明,限制了可靠性、偏见诊断与高效元特征工程。本文系统调研22种现有方法,构建元特征结构化分类体系;采用全局解释技术(决策谓词图)评估选定框架中元模型的特征重要性;并运用局部解释工具(SHAP)分析具体聚类决策。研究揭示了元特征相关性的稳定模式,识别出当前元学习策略中存在的结构性弱点,可能扭曲推荐结果,并为更可解释的自动化机器学习设计提供可操作指导。本研究为提升无监督学习自动化的决策透明性提供了实用基础。
原文摘要 · Abstract (English)
AutoClustering methods aim to automate unsupervised learning tasks, including algorithm selection (AS), hyperparameter optimization (HPO), and pipeline synthesis (PS), by often leveraging meta-learning over dataset meta-features. While these systems often achieve strong performance, their recommendations are often difficult to justify: the influence of dataset meta-features on algorithm and hyperparameter choices is typically not exposed, limiting reliability, bias diagnostics, and efficient meta-feature engineering. This limits reliability and diagnostic insight for further improvements. In this work, we investigate the explainability of the meta-models in AutoClustering. We first review 22 existing methods and organize their meta-features into a structured taxonomy. We then apply a global explainability technique (i.e., Decision Predicate Graphs) to assess feature importance within meta-models from selected frameworks. Finally, we use local explainability tools such as SHAP (SHapley Additive exPlanations) to analyse specific clustering decisions. Our findings highlight consistent patterns in meta-feature relevance, identify structural weaknesses in current meta-learning strategies that can distort recommendations, and provide actionable guidance for more interpretable Automated Machine Learning (AutoML) design. This study therefore offers a practical foundation for increasing decision transparency in unsupervised learning automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。