arXiv:2607.09316cs.CLcs.AI2026-07

用机器学习自动为伏尔泰全集标注主题标签,提升文献检索效率。

Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works

论文配图:Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works
图 1 · 摘自论文原文
  • 将主题索引建模为多标签分类任务,测试不同规模模型表现。
  • 最佳模型(4位量化Mistral)F1达0.67,结果受人工索引主观性影响。
  • 揭示文学修辞特征对自动化处理的挑战,适合数字人文研究者参考。

主题索引——即为文本段落分配结构化概念标签——是大规模文学与历史文献获取的关键,但目前仍主要依赖人工,劳动强度大。本文探索机器学习在自动主题索引中的应用,以伏尔泰全集的两个子语料库(《论各民族风俗与精神》和《关于百科全书的问题》)为测试案例。任务被建模为多标签分类问题,要求模型为一页文本预测专业索引员会添加的主题标签集合。我们对比了从30亿到1200亿参数的多种方法,包括基于编码器的模型及通过低秩适配(LoRA)微调的生成式大语言模型。表现最佳的模型来自Mistral家族,采用4比特量化配置,达到最高F1分数0.67;我们认为该数值为下限,因专业索引本身存在主观性,且模型预测虽与印刷索引不同,仍具语义合理性。此外,我们评估了跨语料泛化能力,并对模型在文学与修辞特征上的表现进行详细定性分析,这些特征尤其难以自动化处理。研究结果对实现大规模文学与历史语料的结构化主题访问具有重要意义。

原文摘要 · Abstract (English)

Thematic indexing -- the practice of assigning structured conceptual labels to sections of text -- is essential to scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labour-intensive process. This paper explores the application of machine learning to automatic thematic indexing, using two substantial sub-corpora of the Complete Works of Voltaire as a test case: the Essai sur les mœurs et l'esprit des nations and the Questions sur l'Encyclopédie. The task is framed as a multi-label classification problem, in which a model must assign the set of index entries that a professional indexer would apply to a given page of text. We compare a range of approaches -- from encoder-based models with classification heads to generative large language models (LLMs) fine-tuned via Low-Rank Adaptation (LoRA) -- spanning model sizes from approximately 3 to 120 billion parameters. Our best-performing model, from the Mistral family in a 4-bit quantised configuration, achieves F1 scores of up to 0.67; we argue that these figures represent lower bounds, given the inherent subjectivity of professional indexing and the frequency with which model predictions prove semantically valid despite diverging from the print index. We further evaluate cross-corpus generalisation and conduct a detailed qualitative analysis of model behaviour on literary and rhetorical features of the source texts that prove particularly resistant to automated treatment. Our findings have implications for the broader challenge of providing structured thematic access to large-scale literary and historical corpora.

主题索引自然语言处理数字人文大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。