用语义新颖度分析2.8万本书,发现叙事结构影响读者数量。
Semantic Novelty at Scale: Narrative Shape Taxonomy and Readership Prediction in 28,606 Books
- 通过句向量与前文中心的余弦距离衡量每段的新颖度。
- 叙事复杂度和结尾/开头新颖度比是预测阅读量最强指标。
- 发现小说与非小说有不同叙事模式,且19世纪书越来越可预测。
本文提出语义新颖度——每段句子嵌入与此前所有段落运行中心的余弦距离——作为大规模叙事结构的信息论度量。基于PG19(1920年前英语文学)中的28,606本书,使用768维SBERT嵌入计算段落级新颖度曲线,并以16段分段聚合近似(PAA)降维。对PAA向量进行Ward链接聚类,识别出八种典型叙事形态,从陡降(快速收敛)到陡升(不断不可预测)。总体波动性(volume)是长度无关的最强读者量预测因子(偏秩相关rho = 0.32),其次为速度(rho = 0.19)与终始比(rho = 0.19)。迂回度原始相关性强(rho = 0.41),但93%与长度相关;控制后偏相关降至0.11,说明文献研究中朴素相关易受长度混淆。体裁强烈限制叙事形态(卡方=2121.6,p < 10⁻²⁴²),小说维持平稳型,非小说信息前置。历史分析显示1840至1910年间书籍逐渐更可预测(T/I比趋势相关r = -0.74,p = 0.037)。SAX分析揭示85%的签名唯一性,表明每本书在语义空间中路径近乎独特。这些发现表明,信息密度动态是独立于情感或主题的根本叙事维度,对读者参与有可测量影响。
原文摘要 · Abstract (English)
I introduce semantic novelty--cosine distance between each paragraph's sentence embedding and the running centroid of all preceding paragraphs--as an information-theoretic measure of narrative structure at corpus scale. Applying it to 28,606 books in PG19 (pre-1920 English literature), I compute paragraph-level novelty curves using 768-dimensional SBERT embeddings, then reduce each to a 16-segment Piecewise Aggregate Approximation (PAA). Ward-linkage clustering on PAA vectors reveals eight canonical narrative shape archetypes, from Steep Descent (rapid convergence) to Steep Ascent (escalating unpredictability). Volume--variance of the novelty trajectory--is the strongest length-independent predictor of readership (partial rho = 0.32), followed by speed (rho = 0.19) and Terminal/Initial ratio (rho = 0.19). Circuitousness shows strong raw correlation (rho = 0.41) but is 93 percent correlated with length; after control, partial rho drops to 0.11--demonstrating that naive correlations in corpus studies can be dominated by length confounds. Genre strongly constrains narrative shape (chi squared = 2121.6, p < 10 to the power negative 242), with fiction maintaining plateau profiles while nonfiction front-loads information. Historical analysis shows books became progressively more predictable between 1840 and 1910 (T/I ratio trend r = negative 0.74, p = 0.037). SAX analysis reveals 85 percent signature uniqueness, suggesting each book traces a nearly unique path through semantic space. These findings demonstrate that information-density dynamics, distinct from sentiment or topic, constitute a fundamental dimension of narrative structure with measurable consequences for reader engagement. Dataset: https://huggingface.co/datasets/wfzimmerman/pg19-semantic-novelty
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。