53种自动化特征工程方法大多难用且缺乏支持,亟需改进可用性。
How Usable is Automated Feature Engineering for Tabular Data?
- 系统评测53种自动特征工程方法的使用体验
- 多数方法无文档、无社区,无法设置时间内存限制
- 适合关注自动化工具实际可用性的研究者与开发者
表格数据在各类机器学习应用中无处不在。每一列代表一个特征,通过组合或转换可生成更具信息量的新特征,这在实现模型最优性能中至关重要。由于手动特征工程成本高、耗时长,已有大量工作致力于自动化。然而,现有自动化特征工程(AutoFE)方法从未被系统评估其对实践者的可用性。我们对53种AutoFE方法进行了调查,发现这些方法普遍难以使用、缺乏文档支持,且无活跃用户社区。此外,所有方法均不支持用户设定时间与内存约束,而我们认为这是实现可用自动化的基本要求。本调研凸显了未来需开发更易用、工程完善的AutoFE方法。
原文摘要 · Abstract (English)
Tabular data, consisting of rows and columns, is omnipresent across various machine learning applications. Each column represents a feature, and features can be combined or transformed to create new, more informative features. Such feature engineering is essential to achieve peak performance in machine learning. Since manual feature engineering is expensive and time-consuming, a substantial effort has been put into automating it. Yet, existing automated feature engineering (AutoFE) methods have never been investigated regarding their usability for practitioners. Thus, we investigated 53 AutoFE methods. We found that these methods are, in general, hard to use, lack documentation, and have no active communities. Furthermore, no method allows users to set time and memory constraints, which we see as a necessity for usable automation. Our survey highlights the need for future work on usable, well-engineered AutoFE methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。