arXiv:2509.08140cs.LGcs.AI2025-09被引 1

用大模型从零数据中挖掘关键特征,预测风投中的稀有高回报项目。

From Limited Data to Rare-event Prediction: LLM-powered Feature Engineering and Multi-model Learning in Venture Capital

  • 用大模型从非结构化数据中提取复杂信号,构建可解释的特征工程
  • 在三个测试集上,精确率比随机基线高出9.8到11.1倍
  • 揭示了创业领域、创始人数量等关键影响因素,适合风投决策者使用

本文提出一种框架,通过整合大语言模型(LLMs)与多模型机器学习架构,实现对罕见但高影响结果的预测。该方法结合黑箱模型的预测力与可解释性需求,利用大模型驱动的特征工程从非结构化数据中提取并合成复杂信号,并在包含XGBoost、随机森林和线性回归的分层集成模型中处理。集成模型先输出连续的成功概率估计,再经阈值化生成二分类的稀有事件预测。该框架应用于风投领域,应对早期数据有限且嘈杂的问题。实证结果显示:在三个独立测试子集上,模型精度达随机基线的9.8至11.1倍。特征敏感性分析进一步揭示:创业公司所属类别贡献15.6%的预测影响力,其次是创始人数量,教育水平和领域经验也表现出小但稳定的正向影响。

原文摘要 · Abstract (English)

This paper presents a framework for predicting rare, high-impact outcomes by integrating large language models (LLMs) with a multi-model machine learning (ML) architecture. The approach combines the predictive strength of black-box models with the interpretability required for reliable decision-making. We use LLM-powered feature engineering to extract and synthesize complex signals from unstructured data, which are then processed within a layered ensemble of models including XGBoost, Random Forest, and Linear Regression. The ensemble first produces a continuous estimate of success likelihood, which is then thresholded to produce a binary rare-event prediction. We apply this framework to the domain of Venture Capital (VC), where investors must evaluate startups with limited and noisy early-stage data. The empirical results show strong performance: the model achieves precision between 9.8X and 11.1X the random classifier baseline in three independent test subsets. Feature sensitivity analysis further reveals interpretable success drivers: the startup's category list accounts for 15.6% of predictive influence, followed by the number of founders, while education level and domain expertise contribute smaller yet consistent effects.

风投预测大模型应用稀有事件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。