arXiv:2509.14448cs.AI2025-09被引 6

首个用于预测初创创始人成败的LLM评测基准,提升投资决策精度。

VCBench: Benchmarking LLMs in Venture Capital

  • 构建9000个匿名创始人数据集,兼顾预测性与隐私保护。
  • 顶尖模型精度超基线6倍,部分超越人类专家表现。
  • 适合关注早期投资、大模型评估与隐私安全的研究者。

SWE-bench和ARC-AGI等基准推动了通用人工智能(AGI)的发展。我们提出VCBench,首个用于预测风险投资中创始人成功概率的基准。该领域信号稀疏、结果不确定,顶级投资者表现也仅略优。初始市场指数精确率为1.9%,Y Combinator表现高出1.7倍,一级机构表现达2.9倍。VCBench提供9000个匿名创始人资料,经标准化处理以保留预测特征并抵御身份泄露,对抗测试显示重识别风险降低超过90%。我们评估了九个前沿大语言模型(LLMs),DeepSeek-V3的精确率超过基线6倍,GPT-4o在F0.5指标上最高,多数模型优于人类基准。该基准作为公开可扩展资源,可通过vcbench.com获取,旨在建立可复现、隐私友好的早期创业预测评估标准。

原文摘要 · Abstract (English)

Benchmarks such as SWE-bench and ARC-AGI demonstrate how shared datasets accelerate progress toward artificial general intelligence (AGI). We introduce VCBench, the first benchmark for predicting founder success in venture capital (VC), a domain where signals are sparse, outcomes are uncertain, and even top investors perform modestly. At inception, the market index achieves a precision of 1.9%. Y Combinator outperforms the index by a factor of 1.7x, while tier-1 firms are 2.9x better. VCBench provides 9,000 anonymized founder profiles, standardized to preserve predictive features while resisting identity leakage, with adversarial tests showing more than 90% reduction in re-identification risk. We evaluate nine state-of-the-art large language models (LLMs). DeepSeek-V3 delivers over six times the baseline precision, GPT-4o achieves the highest F0.5, and most models surpass human benchmarks. Designed as a public and evolving resource available at vcbench.com, VCBench establishes a community-driven standard for reproducible and privacy-preserving evaluation of AGI in early-stage venture forecasting.

大模型评测风险投资隐私保护创始人预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。