用多层级诊断框架评估大模型生成特征的可靠性。
Multi-level Diagnosis and Evaluation for Robust Tabular Feature Engineering with Large Language Models
- 从关键变量、关系和决策边界三方面诊断大模型特征工程
- 不同数据集上表现差异大,优质特征可提升少样本预测10.52%
- 适合关注大模型特征生成可靠性的研究者与工程师
大语言模型在表格数据特征工程中展现出潜力,但其输出稳定性仍存疑。本文提出一个多层级诊断与评估框架,用于衡量大模型在不同领域中特征工程的鲁棒性,重点关注关键变量、变量间关系及预测目标类别的决策边界值。实验表明,大模型的鲁棒性在不同数据集间差异显著,且高质量的生成特征可使少样本预测性能提升最高达10.52%。该工作为评估和提升大模型驱动特征工程的可靠性提供了新方向。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have shown promise in feature engineering for tabular data, but concerns about their reliability persist, especially due to variability in generated outputs. We introduce a multi-level diagnosis and evaluation framework to assess the robustness of LLMs in feature engineering across diverse domains, focusing on the three main factors: key variables, relationships, and decision boundary values for predicting target classes. We demonstrate that the robustness of LLMs varies significantly over different datasets, and that high-quality LLM-generated features can improve few-shot prediction performance by up to 10.52%. This work opens a new direction for assessing and enhancing the reliability of LLM-driven feature engineering in various domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。