用风格分析法揭示中世纪神学著作背后的匿名合作者。
It takes a village to write a book: Mapping anonymous contributions in Stephen Langton's Quaestiones Theologiae
- 通过词频、词性标签和伪词缀分析文本风格,识别作者归属。
- 对比人工撰写与自动提取文本的分析效果,验证方法可靠性。
- 为研究中世纪大学集体写作提供可复用的技术模板。
尽管间接证据表明,早期经院哲学时期基于口头讲授记录(即reportationes)的文献创作并不罕见,但相关直接记载极少。本文设计一项研究,运用文体计量学中的作者归属技术,分析由reportationes汇编而成的《斯蒂芬·朗顿神学问题集》(Stephen Langton's Quaestiones Theologiae),以揭示编辑过程中的多层贡献,并验证有关该文集形成机制的假说。借鉴Camps、Clérice与Pinche(2021)的方法,本文讨论了HTR流水线及基于高频词、词性标记和伪词缀的文体分析实现。该研究将带来两项方法论优势:一是直接比较人工撰写与自动提取数据的表现;二是测试基于Transformer的OCR与自动转录对齐在经院拉丁语语料中的适用性。若成功,该研究将为探索源于中世纪大学的协作性文学创作提供可重复使用的技术范式。
原文摘要 · Abstract (English)
While the indirect evidence suggests that already in the early scholastic period the literary production based on records of oral teaching (so-called reportationes) was not uncommon, there are very few sources commenting on the practice. This paper details the design of a study applying stylometric techniques of authorship attribution to a collection developed from reportationes -- Stephen Langton's Quaestiones Theologiae -- aiming to uncover layers of editorial work and thus validate some hypotheses regarding the collection's formation. Following Camps, Clérice, and Pinche (2021), I discuss the implementation of an HTR pipeline and stylometric analysis based on the most frequent words, POS tags, and pseudo-affixes. The proposed study will offer two methodological gains relevant to computational research on the scholastic tradition: it will directly compare performance on manually composed and automatically extracted data, and it will test the validity of transformer-based OCR and automated transcription alignment for workflows applied to scholastic Latin corpora. If successful, this study will provide an easily reusable template for the exploratory analysis of collaborative literary production stemming from medieval universities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。