arXiv:2501.00138cs.NEcs.AI2025-01

自动构建关联规则挖掘流水线,提升混合数据处理效率

NiaAutoARM: Automated generation and evaluation of Association Rule Mining pipelines

  • 用进化算法自动设计关联规则挖掘全流程
  • 在多个数据集上验证了流水线生成的有效性
  • 适合需要自动化数据分析的科研与工业场景

数值型关联规则挖掘方法能够同时处理数值型和类别型特征,有助于从混合属性数据集中发现关联关系。然而该过程涉及预处理、算法选择、超参数优化及评估指标定义等多个顺序步骤,难以手动完成。本文提出一种新型自动化机器学习方法 NiaAutoARM,基于随机种群启发式算法自动生成完整的关联规则挖掘流水线。文中不仅给出了方法的理论框架,还进行了全面的实验评估,验证了其在多数据集上的有效性。

原文摘要 · Abstract (English)

The Numerical Association Rule Mining paradigm that includes concurrent dealing with numerical and categorical attributes is beneficial for discovering associations from datasets consisting of both features. The process is not considered as easy since it incorporates several processing steps running sequentially that form an entire pipeline, e.g., preprocessing, algorithm selection, hyper-parameter optimization, and the definition of metrics evaluating the quality of the association rule. In this paper, we proposed a novel Automated Machine Learning method, NiaAutoARM, for constructing the full association rule mining pipelines based on stochastic population-based meta-heuristics automatically. Along with the theoretical representation of the proposed method, we also present a comprehensive experimental evaluation of the proposed method.

自动化机器学习关联规则数据挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。