arXiv:2506.20357cs.AI2025-06被引 3

用多种推理方式引导大模型,发现更丰富有用的表格特征。

Tabular Feature Discovery With Reasoning Type Exploration

  • 引入多种推理类型指导大模型生成特征
  • 在59个数据集上提升预测准确率并增加特征多样性
  • 适合需要高质量特征工程的机器学习研究者

表格数据的特征工程仍是机器学习中的关键且挑战性任务。近期,大语言模型(LLMs)被用于利用其海量知识自动生成新特征。然而,现有基于LLM的方法常产生过于简单或重复的特征,部分源于模型自身转换偏见以及生成过程中缺乏结构化推理引导。本文提出一种新方法REFeat,通过引入多种推理类型来引导LLM,以发现更多样、更具信息量的特征。在59个基准数据集上的实验表明,该方法不仅平均预测准确率更高,还生成了更多样且有意义的特征。结果凸显了将丰富推理范式与自适应策略选择融入基于LLM的表格特征发现中的潜力。

原文摘要 · Abstract (English)

Feature engineering for tabular data remains a critical yet challenging step in machine learning. Recently, large language models (LLMs) have been used to automatically generate new features by leveraging their vast knowledge. However, existing LLM-based approaches often produce overly simple or repetitive features, partly due to inherent biases in the transformations the LLM chooses and the lack of structured reasoning guidance during generation. In this paper, we propose a novel method REFeat, which guides an LLM to discover diverse and informative features by leveraging multiple types of reasoning to steer the feature generation process. Experiments on 59 benchmark datasets demonstrate that our approach not only achieves higher predictive accuracy on average, but also discovers more diverse and meaningful features. These results highlight the promise of incorporating rich reasoning paradigms and adaptive strategy selection into LLM-driven feature discovery for tabular data.

特征工程大模型表格数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。