arXiv:2409.16218cs.LGcs.AI2024-09被引 2

PoAC动态适配聚类任务,自动优化分析流程。

Problem-oriented AutoML in Clustering

  • 根据任务目标灵活搭配评估指标与元特征,动态构建聚类方案
  • 在多种数据集上优于现有框架,可视化任务表现更优
  • 无需重新训练,可自动适应不同复杂度数据集

问题导向的聚类自动化机器学习(PoAC)框架提出一种新颖且灵活的聚类自动化方法,克服传统AutoML依赖固定内部聚类有效性指标(CVIs)和静态元特征的局限。与之不同,PoAC建立聚类问题、CVIs与元特征间的动态关联,支持用户按具体任务上下文定制组件。其核心采用基于大规模历史聚类数据集与解决方案的元知识库训练的代理模型,能够推断新聚类流程的质量并合成最优解。相比多数固定评价指标与算法集的AutoML框架,PoAC具备算法无关性,无需额外数据或重训练即可无缝适配不同聚类任务。实验表明,PoAC在多种数据集上超越当前先进框架,尤其在数据可视化等特定任务中表现突出,并能依据数据集复杂度动态调整管道配置。

原文摘要 · Abstract (English)

The Problem-oriented AutoML in Clustering (PoAC) framework introduces a novel, flexible approach to automating clustering tasks by addressing the shortcomings of traditional AutoML solutions. Conventional methods often rely on predefined internal Clustering Validity Indexes (CVIs) and static meta-features, limiting their adaptability and effectiveness across diverse clustering tasks. In contrast, PoAC establishes a dynamic connection between the clustering problem, CVIs, and meta-features, allowing users to customize these components based on the specific context and goals of their task. At its core, PoAC employs a surrogate model trained on a large meta-knowledge base of previous clustering datasets and solutions, enabling it to infer the quality of new clustering pipelines and synthesize optimal solutions for unseen datasets. Unlike many AutoML frameworks that are constrained by fixed evaluation metrics and algorithm sets, PoAC is algorithm-agnostic, adapting seamlessly to different clustering problems without requiring additional data or retraining. Experimental results demonstrate that PoAC not only outperforms state-of-the-art frameworks on a variety of datasets but also excels in specific tasks such as data visualization, and highlight its ability to dynamically adjust pipeline configurations based on dataset complexity.

聚类AutoML自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。