arXiv:2509.10698cs.LGcs.CV2025-09被引 1

用大模型融合结构化与文本数据,提升初创企业成功预测准确率。

CrunchLLM: Multitask LLMs for Structured Business Reasoning and Outcome Prediction

  • 融合公司结构数据与文本描述,通过参数高效微调增强领域适应性。
  • 在Crunchbase数据集上预测准确率超80%,显著优于传统方法和基线模型。
  • 生成可解释推理链条,适合风险投资与政策制定者使用。

预测初创企业是否成功退出(并购或IPO)是创业与创新研究中的关键问题。Crunchbase等数据集包含结构化信息(如融资轮次、行业、投资者网络)和非结构化文本(如公司介绍),但如何有效利用这种异构数据进行预测仍具挑战。传统机器学习方法仅依赖结构化特征,准确率有限;而大语言模型虽具备强大推理能力,却难以直接适配特定商业数据。本文提出 extbf{CrunchLLM},一种针对创业数据的领域自适应大模型框架。该框架将结构化属性与非结构化文本叙事结合,采用参数高效微调与提示优化策略,使基础模型专门化于创业数据。实验显示,CrunchLLM在Crunchbase上的预测准确率超过80%,显著优于传统分类器与基线大模型。此外,该模型还能提供可解释的推理轨迹,增强金融与政策决策者的信任度。本工作展示了通过领域感知微调与结构-文本数据融合,可显著提升创业结果预测建模能力。CrunchLLM不仅提供方法论框架,也是一款面向风投与创新政策的数据驱动工具。

原文摘要 · Abstract (English)

Predicting the success of start-up companies, defined as achieving an exit through acquisition or IPO, is a critical problem in entrepreneurship and innovation research. Datasets such as Crunchbase provide both structured information (e.g., funding rounds, industries, investor networks) and unstructured text (e.g., company descriptions), but effectively leveraging this heterogeneous data for prediction remains challenging. Traditional machine learning approaches often rely only on structured features and achieve moderate accuracy, while large language models (LLMs) offer rich reasoning abilities but struggle to adapt directly to domain-specific business data. We present \textbf{CrunchLLM}, a domain-adapted LLM framework for startup success prediction. CrunchLLM integrates structured company attributes with unstructured textual narratives and applies parameter-efficient fine-tuning strategies alongside prompt optimization to specialize foundation models for entrepreneurship data. Our approach achieves accuracy exceeding 80\% on Crunchbase startup success prediction, significantly outperforming traditional classifiers and baseline LLMs. Beyond predictive performance, CrunchLLM provides interpretable reasoning traces that justify its predictions, enhancing transparency and trustworthiness for financial and policy decision makers. This work demonstrates how adapting LLMs with domain-aware fine-tuning and structured--unstructured data fusion can advance predictive modeling of entrepreneurial outcomes. CrunchLLM contributes a methodological framework and a practical tool for data-driven decision making in venture capital and innovation policy.

大模型创业预测可解释性结构化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。