用多智能体协作让AI自动发现有科学意义的特征。
Knowledge-Informed Automatic Feature Extraction via Collaborative Large Language Model Agents
- 三智能体协同:科学家出思路、提取器生成特征、测试器验证效果。
- 在28个数据集上超越顶尖方法,且能发现如新生物标志物等新假设。
- 结合外部知识库,生成既有效又可解释的特征,适合科研探索者。
机器学习在表格数据上的表现高度依赖高质量的特征工程。尽管大语言模型(LLM)在自动化特征提取(AutoFE)方面展现出潜力,但现有方法受限于单一架构、简单的定量反馈以及缺乏对领域知识的系统性整合。本文提出Rogue One,一种基于LLM的多智能体框架,用于知识引导的自动化特征提取。该框架采用去中心化的三个专用智能体——科学家、提取器和测试器——通过迭代协作实现特征的发现、生成与验证。关键在于,系统摒弃了简单的准确率评分,引入丰富的定性反馈机制和“泛滥-修剪”策略,动态平衡特征探索与利用。通过集成检索增强(RAG)系统主动融入外部知识,Rogue One生成的特征不仅统计性能强,且语义明确、可解释。我们在涵盖19个分类和9个回归任务的广泛数据集上验证了其显著优于当前最优方法的表现。此外,定性分析显示该系统能揭示新且可检验的科学假设,例如在心肌数据集中识别出潜在的新生物标志物,凸显其作为科研发现工具的价值。
原文摘要 · Abstract (English)
The performance of machine learning models on tabular data is critically dependent on high-quality feature engineering. While Large Language Models (LLMs) have shown promise in automating feature extraction (AutoFE), existing methods are often limited by monolithic LLM architectures, simplistic quantitative feedback, and a failure to systematically integrate external domain knowledge. This paper introduces Rogue One, a novel, LLM-based multi-agent framework for knowledge-informed automatic feature extraction. Rogue One operationalizes a decentralized system of three specialized agents-Scientist, Extractor, and Tester-that collaborate iteratively to discover, generate, and validate predictive features. Crucially, the framework moves beyond primitive accuracy scores by introducing a rich, qualitative feedback mechanism and a "flooding-pruning" strategy, allowing it to dynamically balance feature exploration and exploitation. By actively incorporating external knowledge via an integrated retrieval-augmented (RAG) system, Rogue One generates features that are not only statistically powerful but also semantically meaningful and interpretable. We demonstrate that Rogue One significantly outperforms state-of-the-art methods on a comprehensive suite of 19 classification and 9 regression datasets. Furthermore, we show qualitatively that the system surfaces novel, testable hypotheses, such as identifying a new potential biomarker in the myocardial dataset, underscoring its utility as a tool for scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。