用大模型快速生成真实实体的表格数据,验证关于作家童年经历等假设。
Simulating Tabular Datasets through LLMs to Rapidly Explore Hypotheses about Real-World Entities
- 用大模型估算具体人物、公司等实体的属性值,生成可分析的表格数据。
- 模型规模越大,对实体属性的估计越准确,支持定量分析。
- 适合想快速验证社会科学研究假设的研究者使用。
许多作家是否拥有更糟糕的童年?尽管许多作家的生平资料已知,但量化验证此类定性假设需大量人工工作,如筛选大量传记与访谈,并迭代寻找能反映质性关注点的量化特征。本文探索通过(1)利用大模型估算具体实体(如特定人物、公司、书籍、动物种类、国家)的属性;(2)应用现成分析方法揭示这些属性间的潜在关系(如线性回归);以及(3)进一步自动化,让大模型建议可用于支撑特定定性假设的量化属性(如在示例中提出“童年负面事件数量”)。目标是实现人机协作,加速假设筛选。实验表明,大模型可在多个领域有效估算具体实体的表格数据,且性能随模型规模提升。初步实验还展示大模型将定性假设映射为可估算的具体变量的潜力。结论是,大模型有望揭示其训练数据中潜藏的科学有趣模式。
原文摘要 · Abstract (English)
Do horror writers have worse childhoods than other writers? Though biographical details are known about many writers, quantitatively exploring such a qualitative hypothesis requires significant human effort, e.g. to sift through many biographies and interviews of writers and to iteratively search for quantitative features that reflect what is qualitatively of interest. This paper explores the potential to quickly prototype these kinds of hypotheses through (1) applying LLMs to estimate properties of concrete entities like specific people, companies, books, kinds of animals, and countries; (2) performing off-the-shelf analysis methods to reveal possible relationships among such properties (e.g. linear regression); and towards further automation, (3) applying LLMs to suggest the quantitative properties themselves that could help ground a particular qualitative hypothesis (e.g. number of adverse childhood events, in the context of the running example). The hope is to allow sifting through hypotheses more quickly through collaboration between human and machine. Our experiments highlight that indeed, LLMs can serve as useful estimators of tabular data about specific entities across a range of domains, and that such estimations improve with model scale. Further, initial experiments demonstrate the potential of LLMs to map a qualitative hypothesis of interest to relevant concrete variables that the LLM can then estimate. The conclusion is that LLMs offer intriguing potential to help illuminate scientifically interesting patterns latent within the internet-scale data they are trained upon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。