arXiv:2608.00045cs.CLecon.GN2026-08

仅靠创业描述文本就能预测早期成功,无需财务或人力数据。

Predicting Startup Exit from Textual Descriptors - A Computational Linguistics Framework

  • 用850个文本特征构建创业叙事映射,提取语言风格信号。
  • 纯文本预测F1达0.30,最优模型LightGBM达0.48。
  • 提出可量化的'煽动性评分',适合风投快速评估初创前景。

本研究发现,仅依靠文本描述即可预测早期创业公司成功(以退出为标准),无需依赖上下文、财务或人力资本变量。基于20年跨度、涵盖7,419家初创企业的风投整理数据集,研究提取文本框架变量,并通过创业叙事映射生成850个特征。对数据子集与向量嵌入进行显著性检验后,开展六种监督学习模型实验。LightGBM表现最佳(F1 = 0.48),而纯文本特征已实现F1 = 0.30,证实创始人叙述的独立预测价值。特征分析显示,适度使用夸张标记(如形容词、行话、流行词)与更高退出概率相关;过度冗长的陈述或名称则降低预测效果。研究还引入可量化的‘煽动性评分’用于风投申请,证明在信息高度不对称条件下,创业表述可提供可度量的成功信号。

原文摘要 · Abstract (English)

This study shows that textual descriptors alone can predict early-stage startup success, defined as Exit, without relying on contextual, financial, or human capital variables. Using venture capital-curated datasets covering 7,419 startups over 20 years, the research isolates text-based framing variables and engineers 850 features through startup narrative mapping. Data subsets and vector embeddings are evaluated for statistical significance, followed by supervised machine learning experiments across six models. LightGBM achieved the highest predictive performance (F1 = 0.48), while textual descriptors alone achieved F1 = 0.30, confirming the standalone predictive value of founder narratives. Feature analysis shows that optimized densities of hyping markers, including adjectives, jargon, and buzzwords, are associated with higher Exit probability, whereas excessive statement or name length reduces it. The study also introduces a quantifiable Hyping Score for venture capital applications, demonstrating that startup framing provides measurable signals for predicting Exit under conditions of high information asymmetry.

文本预测创业评估自然语言处理风投

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。