arXiv:2509.02605cs.MAcs.AI2025-09

用AI生成创业者模拟人物,对比真实访谈,验证其在创业验证研究中的有效性。

Synthetic Founders: AI-Generated Social Simulations for Startup Validation Research in Computational Social Science

  • 用大语言模型生成虚拟创业者和投资人,与真人访谈对比分析。
  • 发现AI模拟能复现真实用户对效率和信任的担忧,但忽略历史创伤等深层问题。
  • 适合做创业早期假设探索,补充真实调研的盲区。

我们开展了一项对比性对接实验,将真人初创企业创始人访谈数据与大语言模型(LLM)驱动的虚拟角色数据进行比对,评估人工智能模拟在真实性、差异性和盲点方面的表现。15位早期阶段创业者接受了关于人工智能赋能验证的希望与顾虑访谈,同一方法也应用于由AI生成的创业者与投资人角色。结构化主题合成揭示四类结果:(1)一致主题——承诺型需求信号、黑箱信任障碍与效率提升在两组数据中均被强调;(2)部分重叠——真人关注异常值被平均化及真实客户验证压力,而虚拟角色则突出非理性盲点,并将AI视为心理缓冲;(3)仅人类拥有主题——早期客户互动的亲缘价值与对远大市场的怀疑;(4)仅虚拟角色呈现主题——放大假阳性与创伤盲点,即AI可能因忽视负面历史经验而高估采纳潜力。我们认为,这种对比框架表明,基于LLM的角色构成一种混合社会模拟:语言表达更丰富灵活,但缺乏真实经历与关系后果。它们不替代实证研究,而是作为补充性模拟类型,可拓展假设空间、加速探索性验证,并明确计算社会科学研究中认知真实性的边界。

原文摘要 · Abstract (English)

We present a comparative docking experiment that aligns human-subject interview data with large language model (LLM)-driven synthetic personas to evaluate fidelity, divergence, and blind spots in AI-enabled simulation. Fifteen early-stage startup founders were interviewed about their hopes and concerns regarding AI-powered validation, and the same protocol was replicated with AI-generated founder and investor personas. A structured thematic synthesis revealed four categories of outcomes: (1) Convergent themes - commitment-based demand signals, black-box trust barriers, and efficiency gains were consistently emphasized across both datasets; (2) Partial overlaps - founders worried about outliers being averaged away and the stress of real customer validation, while synthetic personas highlighted irrational blind spots and framed AI as a psychological buffer; (3) Human-only themes - relational and advocacy value from early customer engagement and skepticism toward moonshot markets; and (4) Synthetic-only themes - amplified false positives and trauma blind spots, where AI may overstate adoption potential by missing negative historical experiences. We interpret this comparative framework as evidence that LLM-driven personas constitute a form of hybrid social simulation: more linguistically expressive and adaptable than traditional rule-based agents, yet bounded by the absence of lived history and relational consequence. Rather than replacing empirical studies, we argue they function as a complementary simulation category - capable of extending hypothesis space, accelerating exploratory validation, and clarifying the boundaries of cognitive realism in computational social science.

社会模拟AI生成创业验证大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。