arXiv:2605.02974q-fin.PRcs.LG2026-05

用 Product Hunt 发布数据预测初创公司能否拿到 A 轮融资

PHBench: A Benchmark for Predicting Startup Series A Funding from Product Hunt Launch Signals

论文配图:PHBench: A Benchmark for Predicting Startup Series A Funding from Product Hunt Launch Signals
图 1 · 摘自论文原文
  • 构建包含 6.7 万条发布记录的基准数据集,关联融资信息
  • 模型在测试集上实现 F0.5=0.097,比随机预测高 4.7 倍
  • 公开数据、代码与排行榜,支持复现和对比评估

Product Hunt 上的结构化发布信号对 A 轮融资结果具有统计显著的预测能力。我们基于 2019–2025 年间 67,292 条已发布内容,通过确定性域名匹配链接 Crunchbase 融资记录,识别出 528 次在发布后 18 个月内完成的验证型 A 轮融资(正样本率 0.78%)。最佳模型为三组件集成(ENS_avg、ENS_ISO、XGB),经验证 F0.5 = 0.097,AP = 0.037(95% 置信区间:0.024–0.072;较随机预测提升 4.7 倍);配对自助法证实其显著优于逻辑回归基线(AP 差值 +0.013,95% 置信区间 [0.004, 0.039],p < 0.001;F0.5 差值 +0.056,95% 置信区间 [0.006, 0.122],p = 0.016)。验证集指标(F0.5 = 0.284,AP = 0.126)反映对 53 个正例的前 144 模型选择偏差,仅用于基准可复现性说明。我们还评估了三种零样本 Gemini 模型(Gemini 2.5 Flash、Gemini 3 Flash、Gemini 3.1 Pro)在匿名数值设定下的表现,最优者(Gemini 3 Flash)AP = 0.034,低于逻辑回归基线的 0.044。值得注意的是,性能最强的 Gemini 3.1 Pro(AP = 0.023)反而表现最差,这一反常现象需进一步探究跨模型与提示策略的影响。机器学习与大模型均表现出与 2020–2021 年融资热潮及后续收缩同步的时间衰减趋势,验证数据集捕捉了真实的市场结构而非噪声。PHBench 提供可复现框架,包含公开训练/验证/盲测划分、61 个工程特征、五项评估指标体系及公开排行榜(https://phbench.com),所有代码、基线模型与匿名数据分片均已开源。

原文摘要 · Abstract (English)

Structured launch signals on Product Hunt contain statistically significant predictive information for Series A funding outcomes. We construct PHBench from 67,292 featured Product Hunt posts spanning 2019-2025, linked to Crunchbase funding records via deterministic domain matching, identifying 528 verified Series A raises within 18 months of launch (positive rate: 0.78%). Our best-performing model, a three-component ensemble (ENS_avg, ENS_ISO, XGB) selected by validation F0.5, achieves F0.5 = 0.097 and AP = 0.037 (95% CI: 0.024-0.072; 4.7x lift over random) on the private held-out test set (103 positives). A paired bootstrap confirms a statistically credible advantage over the logistic regression baseline (AP delta: +0.013, 95% CI: [0.004, 0.039], p < 0.001; F0.5 delta: +0.056, 95% CI: [0.006, 0.122], p = 0.016). Validation-set metrics (F0.5 = 0.284, AP = 0.126) reflect best-of-144 selection bias on 53 positives and are reported for benchmark reproducibility only. We further evaluate three zero-shot Gemini models (Gemini 2.5 Flash, Gemini 3 Flash, and Gemini 3.1 Pro) in an anonymized numerical setting. The best LLM achieves AP = 0.034 (Gemini 3 Flash), below the LR baseline AP of 0.044. Notably, the most capable Gemini variant (Gemini 3.1 Pro, AP = 0.023) performs worst -- an unexpected pattern that warrants further investigation across providers and prompting strategies. Both ML and LLM models show the same temporal performance decay tracking the 2020-2021 funding boom and subsequent contraction, confirming the dataset captures genuine market structure rather than noise. PHBench provides a reproducible framework comprising public training, validation, and blind test splits; 61 engineered features; a five-metric evaluation harness; and a public leaderboard at https://phbench.com. All code, baseline models, and anonymized dataset splits are publicly available.

创业融资预测模型数据基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。