arXiv:2503.20394cs.LGcs.AI2025-03中稿 · ICDE 2025被引 6

用智能探索策略加速特征变换,提升自动化数据工程效率。

FastFT: Accelerating Reinforced Feature Transformation via Advanced Exploration Strategies

  • 分离评估与下游任务,用预测器快速打分
  • 引入新颖性奖励,解决有效变换反馈稀疏问题
  • 优先记忆关键经验,适合大规模特征工程场景

特征变换对以数据为中心的经典机器学习至关重要,旨在生成特征组合以提升下游任务性能。现有方法如人工设计、迭代反馈和探索生成策略虽能减少人工干预,但仍面临三大挑战:(1) 依赖下游任务指标评估,耗时长,尤其在大数据集上;(2) 随机探索后难以保证特征组合多样性;(3) 罕见的显著变换导致有价值反馈稀疏,阻碍学习进程或降低效果。针对这些问题,本文提出FastFT框架,采用三项先进策略:首先通过性能预测器将特征变换评估与生成数据集结果解耦,实现快速评估;其次设计新颖性评估方法,将其融入奖励函数,加速模型对有效变换的探索;此外,结合新颖性与性能构建优先级记忆缓冲区,确保关键经验被有效重访。大量实验验证了该框架在性能、效率与可追溯性方面的优势,展现出在复杂特征变换任务中的卓越能力。

原文摘要 · Abstract (English)

Feature Transformation is crucial for classic machine learning that aims to generate feature combinations to enhance the performance of downstream tasks from a data-centric perspective. Current methodologies, such as manual expert-driven processes, iterative-feedback techniques, and exploration-generative tactics, have shown promise in automating such data engineering workflow by minimizing human involvement. However, three challenges remain in those frameworks: (1) It predominantly depends on downstream task performance metrics, as assessment is time-consuming, especially for large datasets. (2) The diversity of feature combinations will hardly be guaranteed after random exploration ends. (3) Rare significant transformations lead to sparse valuable feedback that hinders the learning processes or leads to less effective results. In response to these challenges, we introduce FastFT, an innovative framework that leverages a trio of advanced strategies.We first decouple the feature transformation evaluation from the outcomes of the generated datasets via the performance predictor. To address the issue of reward sparsity, we developed a method to evaluate the novelty of generated transformation sequences. Incorporating this novelty into the reward function accelerates the model's exploration of effective transformations, thereby improving the search productivity. Additionally, we combine novelty and performance to create a prioritized memory buffer, ensuring that essential experiences are effectively revisited during exploration. Our extensive experimental evaluations validate the performance, efficiency, and traceability of our proposed framework, showcasing its superiority in handling complex feature transformation tasks.

特征工程强化学习自动化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。