系统梳理表格数据增强方法,提升模型性能。
Towards Data-Centric AI: A Comprehensive Survey of Traditional, Reinforcement, and Generative Approaches for Tabular Data Transformation
- 从特征选择与生成两方面优化表格数据表示
- 涵盖传统、强化学习与生成式方法的全面分析
- 适合关注数据质量与特征工程的研究者
表格数据是金融、医疗、营销等领域广泛应用的数据格式。在以数据为中心的人工智能时代,提升数据质量和表征对模型性能至关重要,尤其在表格数据应用中。本文系统综述了表格数据为中心的AI关键技术,重点聚焦特征选择与特征生成,前者用于识别并保留关键属性,后者通过构造新特征捕捉复杂模式。通过对近年进展、实际应用及各类方法优缺点的分析,提供完整方法论概览,并指出当前挑战与未来方向,推动该领域持续创新。
原文摘要 · Abstract (English)
Tabular data is one of the most widely used formats across industries, driving critical applications in areas such as finance, healthcare, and marketing. In the era of data-centric AI, improving data quality and representation has become essential for enhancing model performance, particularly in applications centered around tabular data. This survey examines the key aspects of tabular data-centric AI, emphasizing feature selection and feature generation as essential techniques for data space refinement. We provide a systematic review of feature selection methods, which identify and retain the most relevant data attributes, and feature generation approaches, which create new features to simplify the capture of complex data patterns. This survey offers a comprehensive overview of current methodologies through an analysis of recent advancements, practical applications, and the strengths and limitations of these techniques. Finally, we outline open challenges and suggest future perspectives to inspire continued innovation in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。