构建百万级人物社交网络,揭示19世纪以来全球文学中的社会关系图谱
A City of Millions: Mapping Literary Social Networks At Scale
- 用语言模型自动提取跨语种小说非虚构作品中的社交关系
- 涵盖251万个人、280万对关系,覆盖1800-1999年58种语言
- 适合人文社科研究者分析历史社会认知模式
我们发布了从多语种小说与非虚构叙事中提取的70,509个高质量社交网络。此外还提供约30,000篇文本的元数据(73%非虚构,27%小说),时间跨度为1800至1999年,涉及58种语言。该数据集以前所未有的规模呈现了历史社会世界的信息,包含2,510,021名个体和2,805,482对人际关系,每对关系均标注亲密度与关系类型。我们通过将以往人工标注流程转化为语言模型提示,实现大规模一致性抽取。该数据集为人文与社会科学提供了独特资源,支持对社会现实认知模型的研究。
原文摘要 · Abstract (English)
We release 70,509 high-quality social networks extracted from multilingual fiction and nonfiction narratives. We additionally provide metadata for $\sim$30,000 of these texts (73\% nonfiction and 27\% fiction) written between 1800 and 1999 in 58 languages. This dataset provides information on historical social worlds at an unprecedented scale, including data for 2,510,021 individuals in 2,805,482 pair-wise relationships annotated for affinity and relationship type. We achieve this scale by automating previously manual methods of extracting social networks; specifically, we adapt an existing annotation task as a language model prompt, ensuring consistency at scale with the use of structured output. This dataset serves as a unique resource for humanities and social science research by providing data on cognitive models of social realities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。