arXiv:2603.01791cs.CLcs.IR2026-03

分析8万本书的语义新颖性轨迹,发现现代小说更复杂但文学价值不相关。

Semantic Novelty Trajectories in 80,000 Books: A Cross-Corpus Embedding Analysis

  • 用嵌入向量追踪每段文字的语义新颖度变化
  • 现代书籍新颖度高10%,路径更曲折,但经典叙事更常见于旧书
  • 新颖性与读者评分无关,适合对叙事结构感兴趣的学者

本文在语料库尺度上应用Schmidhuber的压缩进度趣味性理论,分析了跨越两个世纪、超过8万本英语出版物的语义新颖性轨迹。通过句子变换器段落嵌入与运行中心点新颖性度量,比较了28,730本1920年前的Project Gutenberg图书(PG19)与52,796本现代英文书籍(Books3,约1990–2010)。主要发现有四点:第一,现代书籍段落级新颖性平均高出10%(0.503 vs. 0.459);第二,轨迹迂回度(嵌入空间中累计路径长度与净位移之比)在现代文献中几乎翻倍(+67%);第三,新颖性递减趋于稳定语义的收敛叙事曲线,在1920年前文献中出现频率是现代书的2.3倍;第四,新颖性与读者评分高度无关(r = -0.002),表明该理论中的趣味性在结构上独立于文学评价。基于PAA-16表示聚类段落轨迹,识别出八种叙事形态原型,其分布随时代显著变化。所有分析代码及交互式探索工具已公开于https://bigfivekiller.online/novelty_hub。

原文摘要 · Abstract (English)

I apply Schmidhuber's compression progress theory of interestingness at corpus scale, analyzing semantic novelty trajectories in more than 80,000 books spanning two centuries of English-language publishing. Using sentence-transformer paragraph embeddings and a running-centroid novelty measure, I compare 28,730 pre-1920 Project Gutenberg books (PG19) against 52,796 modern English books (Books3, approximately 1990-2010). The principal findings are fourfold. First, mean paragraph-level novelty is roughly 10% higher in modern books (0.503 vs. 0.459). Second, trajectory circuitousness -- the ratio of cumulative path length to net displacement in embedding space -- nearly doubles in the modern corpus (+67%). Third, convergent narrative curves, in which novelty declines toward a settled semantic register, are 2.3x more common in pre-1920 literature. Fourth, novelty is orthogonal to reader quality ratings (r = -0.002), suggesting that interestingness in Schmidhuber's sense is structurally independent of perceived literary merit. Clustering paragraph-level trajectories via PAA-16 representations reveals eight distinct narrative-shape archetypes whose distribution shifts substantially between eras. All analysis code and an interactive exploration toolkit are publicly available at https://bigfivekiller.online/novelty_hub.

语义分析叙事结构文本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。