用大模型推理自动生成机器学习特征,省去人工设计规则。
Applications of Large Language Model Reasoning in Feature Generation
- 用链式思考等四种推理方法自动发现有效特征生成规则。
- 在医疗、金融等领域提升数据利用效率,支持文本生成与实体提取。
- 适合想自动化特征工程的研究者和工程师,尤其关注低资源场景。
大型语言模型(LLMs)凭借其先进的推理能力,彻底改变了自然语言处理。本文探讨了LLM推理技术与特征生成在机器学习任务中的融合。研究分析了四种关键推理方法:链式思考(Chain of Thought)、思维树(Tree of Thoughts)、检索增强生成(Retrieval-Augmented Generation)和思想空间探索(Thought Space Exploration)。结果表明,这些方法可在无需手动定义搜索空间的情况下,识别出有效的特征生成规则。论文对金融、医疗、文本分析等多个领域的基于LLM的特征生成方法进行了分类。在医疗领域,LLMs可从临床记录和放射报告中提取关键信息,实现更高效的数据利用;在金融领域,则支持复杂文档的文本生成、摘要和实体抽取。文章还分析了评估特征质量与下游性能的方法,重点关注OCTree的决策树推理方法,该方法通过语言反馈实现迭代优化。当前挑战包括幻觉、计算效率和领域适应性。截至2025年3月,新兴方法包括推理时计算扩展、强化学习以及带模型蒸馏的监督微调。未来方向包括多模态特征生成、自改进系统和神经符号方法。本文全面概述了这一新兴领域,展示了通过语言模型推理实现特征工程自动化与增强的巨大潜力。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have revolutionized natural language processing through their state of art reasoning capabilities. This paper explores the convergence of LLM reasoning techniques and feature generation for machine learning tasks. We examine four key reasoning approaches: Chain of Thought, Tree of Thoughts, Retrieval-Augmented Generation, and Thought Space Exploration. Our analysis reveals how these approaches can be used to identify effective feature generation rules without having to manually specify search spaces. The paper categorizes LLM-based feature generation methods across various domains including finance, healthcare, and text analytics. LLMs can extract key information from clinical notes and radiology reports in healthcare, by enabling more efficient data utilization. In finance, LLMs facilitate text generation, summarization, and entity extraction from complex documents. We analyze evaluation methodologies for assessing feature quality and downstream performance, with particular attention to OCTree's decision tree reasoning approach that provides language-based feedback for iterative improvements. Current challenges include hallucination, computational efficiency, and domain adaptation. As of March 2025, emerging approaches include inference-time compute scaling, reinforcement learning, and supervised fine-tuning with model distillation. Future directions point toward multimodal feature generation, self-improving systems, and neuro-symbolic approaches. This paper provides a detailed overview of an emerging field that promises to automate and enhance feature engineering through language model reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。