arXiv:2603.00059cs.CYcs.AI2026-03综述被引 1

五款主流大模型生成的调研数据无法还原人类真实洞察,仅能复刻共识。

Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data

  • 用五款大模型模拟420名硅谷程序员调研,生成合成数据
  • 合成数据整体趋同,与真实人类数据差异显著,缺乏反直觉发现
  • 适合用于识别社会共识,不适合替代真实调研

本文比较了420名硅谷程序员的真实调研数据与五款领先生成式AI大模型(ChatGPT Thinking 5 Pro、Claude Sonnet 4.5 Pro + Claude CoWork 1.123、Gemini Advanced 2.5 Pro、Incredible 1.0、DeepSeek 3.2)生成的合成调研数据。结果表明,尽管合成数据在技术上合理且趋于一致,但未能捕捉真实调研中具反直觉价值的洞见;所有模型生成的数据彼此趋同,真实数据反而成为异常值。研究发现,当前主流大模型虽可高效复制和扩展调研,但仅强化共识而非揭示新知。若未来使用合成受访者,需建立可重复的验证协议与使用标准。结论认为,合成调研数据不应取代严谨调研方法,而应作为田野调查前后的辅助工具,用于识别社会规范、普遍认知与群体预期。

原文摘要 · Abstract (English)

How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2. Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier. Our key finding is that while leading LLMs are increasingly being used to scale, replicate and replace human survey responses in research, these advances only show an increased capacity to parrot conventional wisdom in harmony with each other rather than revealing novel findings. If synthetic respondents are used in future research, we need more replicable validation protocols and reporting standards for when and where synthetic survey data can be used responsibly, a gap that this paper fills. Our results suggest that synthetic survey responses cannot meaningfully model real human social beliefs within organizations, particularly in contexts lacking previously documented evidence. We conclude that synthetic survey-based research should be cast not as a substitute for rigorous survey methods, but as an increasingly reliable pre- or post-fieldwork instrument for identifying societal assumptions, conventional wisdoms, and other expectations about research populations.

大模型调研数据合成数据组织研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。