构建53万本数字书籍目录,助力文化研究与社会计算。
MajinBook: An open catalogue of digitally mediated world literature
- 整合影子图书馆与Goodreads数据,构建高精度数字书目
- 覆盖三百年英语图书,含出版时间、类型、评分等元数据
- 开源数据并讨论合规性,支持多语言研究
本文介绍MajinBook,一个开放书目库,旨在促进影子图书馆(如Library Genesis和Z-Library)在计算社会科学与文化分析中的应用。通过将这些大规模众包档案的元数据与Goodreads的结构化书目数据关联,我们构建了一个包含超过539,000条英文数字书籍记录的高精度语料库。该语料库跨越三个世纪,反映当代选书偏差,包含首次出版日期、类别及评分、评论等流行度指标。方法上优先采用原生数字EPUB文件以保障机器可读性,缓解传统语料库(如HathiTrust)的偏差问题,并包含法语、德语和西班牙语的次级数据集。我们评估了数据关联策略的准确性,公开所有底层数据,并讨论该项目在欧盟与美国文本与数据挖掘研究框架下的法律可允许性。
原文摘要 · Abstract (English)
This data paper introduces MajinBook, an open catalogue designed to facilitate the use of shadow libraries-such as Library Genesis and Z-Library-for computational social science and cultural analytics. By linking metadata from these vast, crowd-sourced archives with structured bibliographic data from Goodreads, we create a high-precision corpus of over 539,000 references to digitally mediated English-language books. Spanning three centuries and reflecting a contemporary selection bias, these entries are enriched with first publication dates, genres, and popularity metrics like ratings and reviews. Our methodology prioritises natively digital EPUB files to ensure machine-readable quality, while addressing biases in traditional corpora like HathiTrust, and includes secondary datasets for French, German, and Spanish. We evaluate the linkage strategy for accuracy, release all underlying data openly, and discuss the project's legal permissibility under EU and US frameworks for text and data mining in research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。