从强化学习与生成模型视角,系统梳理表格数据特征优化方法。
A Survey on Data-Centric AI: Tabular Learning from Reinforcement Learning and Generative AI Perspective
- 用强化学习与生成模型自动优化特征选择与生成
- 总结现有方法在真实场景中的表现与局限性
- 适合关注数据驱动智能建模的研究者与工程师
表格数据是生物信息学、医疗健康和市场营销等领域最常用的数据格式之一。随着人工智能向以数据为中心的范式演进,提升数据质量成为增强表格数据驱动应用模型性能的关键。本综述聚焦于数据驱动的表格数据优化,重点探讨基于强化学习(RL)和生成模型的特征选择与特征生成技术,作为精炼数据空间的基础手段。特征选择旨在识别并保留最具信息量的属性,而特征生成则通过构造新特征以更好捕捉复杂数据模式。我们系统回顾了现有的表格数据生成方法,分析其最新进展、实际应用场景及各自的优势与不足。本综述强调了基于强化学习与生成技术在特征工程自动化与智能化方面的贡献。最后,总结当前挑战并讨论未来研究方向,旨在为该领域持续创新提供洞见。
原文摘要 · Abstract (English)
Tabular data is one of the most widely used data formats across various domains such as bioinformatics, healthcare, and marketing. As artificial intelligence moves towards a data-centric perspective, improving data quality is essential for enhancing model performance in tabular data-driven applications. This survey focuses on data-driven tabular data optimization, specifically exploring reinforcement learning (RL) and generative approaches for feature selection and feature generation as fundamental techniques for refining data spaces. Feature selection aims to identify and retain the most informative attributes, while feature generation constructs new features to better capture complex data patterns. We systematically review existing generative methods for tabular data engineering, analyzing their latest advancements, real-world applications, and respective strengths and limitations. This survey emphasizes how RL-based and generative techniques contribute to the automation and intelligence of feature engineering. Finally, we summarize the existing challenges and discuss future research directions, aiming to provide insights that drive continued innovation in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。