用SHAP值引导特征生成,无元数据也能高效自动建模。
SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE

- 用SHAP值替代元数据,动态指导特征组合生成
- 生成特征重复率从37.2%降至6.8%,平均仅需5.4个特征
- 适合缺乏标签信息的工业级自动化特征工程场景
近期研究利用大语言模型(LLM)通过语义描述和轨迹提示增强自动化特征工程(AutoFE),但在长时程优化中面临两大挑战:(1)实际场景常缺失语义元数据;(2)轨迹累积易超出上下文长度,不累积则生成过程不稳定,易陷入局部最优且特征重复率高。为此,我们提出一种无需元数据的可扩展常量上下文优化框架SIGMA,其结合SHAP值提供任务感知信号以指导分组特征生成,并引入暴露特征隐式轨迹(EXIT)机制,使提示中的已暴露特征隐含表示优化路径。实验表明,SIGMA在几乎恒定提示长度下达到与当前最先进(SOTA)LLM基线相当的性能。值得注意的是,EXIT将特征重复率从37.2%显著降低至6.8%,同时平均仅需5.4个特征即匹配传统SOTA表现,展现出显著的特征利用效率提升。
原文摘要 · Abstract (English)
Recent research has leveraged Large Language Models (LLMs) to enhance Automated Feature Engineering (AutoFE) through semantic descriptions and trajectory-based prompting. However, there exist two challenges that limit their applicability and scalability in long-horizon optimization: (1) semantic metadata is unavailable in many practical settings, and (2) trajectory accumulation increases the risk of exceeding the context window, while without it, the generation process can become unstable, leading to becoming stuck in the local optima and a high duplicate rate of generated features. To this end, we propose a SHAP-enhanced Implicit-trajectory Generation for Metadata-free AutoFE (SIGMA), a scalable constant-context optimization framework. SIGMA leverages SHAP values to provide task-aware signals for guiding group feature generation instead of semantic information. In addition, we adopt an EXposed-feature Implicit Trajectory (EXIT) approach, where the exposed features in the prompt implicitly represent the trajectory. Empirical results demonstrate that SIGMA achieves performance comparable to the state-of-the-art (SOTA) LLM baselines with a nearly constant prompt length. Notably, EXIT significantly reduces the duplicate ratio of generated features from 37.2% to 6.8%. At the same time, SIGMA matches traditional SOTA performance with only 5.4 features on average, demonstrating substantial efficiency gains in feature utilization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。