通过共现模式发现文本的结构功能,而非仅主题内容。
From Topic to Transition Structure: Unsupervised Concept Discovery at Corpus Scale via Predictive Associative Memory
- 用共现对训练小模型,压缩后捕捉文本的重复过渡结构。
- 在k=100时,每类平均包含4508本书,揭示跨作品的文学模式。
- 适合研究文学结构、语言风格或无需标注的大规模文本分析者。
嵌入模型按语义内容分组文本,即‘文本讲什么’。我们发现文本内部的时间共现可揭示另一类结构:重复出现的过渡结构概念,即‘文本做什么’。在9,766部古腾堡计划文本(2496万段落)中,基于3.73亿个共现对,训练了一个2940万参数的对比模型,将预训练嵌入映射到关联空间,使具有相似过渡结构的段落聚集在一起。在容量受限(42.75%准确率)条件下,模型必须压缩重复模式而非记忆个体共现。在六种粒度(k=50至k=2000)下聚类生成多分辨率概念图,涵盖从‘直接对抗’‘抒情冥想’到‘水手方言’‘法庭质询’等精确场景模板。k=100时,每类平均含4508本书(共9766本),验证了全语料范围的规律性。与嵌入相似性聚类对比显示,原始嵌入按主题分组,而关联空间聚类则按功能、语体和文学传统分组。未见小说可直接分配至已有集群,关联模型将每部小说浓缩至少数连贯集群,而原始嵌入分配几乎饱和所有集群。验证控制排除了位置、长度和书籍集中度干扰。该方法将预测性关联记忆(PAM)从情景回忆拓展至概念形成:当PAM用于特定关联回忆时,多周期对比训练在压缩约束下提取可迁移的结构性模式,同一框架在不同范式下产生质变行为。
原文摘要 · Abstract (English)
Embedding models group text by semantic content, what text is about. We show that temporal co-occurrence within texts discovers a different kind of structure: recurrent transition-structure concepts or what text does. We train a 29.4M-parameter contrastive model on 373 million co-occurrence pairs from 9,766 Project Gutenberg texts (24.96 million passages), mapping pre-trained embeddings into an association space where passages with similar transition structure cluster together. Under capacity constraint (42.75% accuracy), the model must compress across recurring patterns rather than memorise individual co-occurrences. Clustering at six granularities (k=50 to k=2,000) produces a multi-resolution concept map; from broad modes like "direct confrontation" and "lyrical meditation" to precise registers and scene templates like "sailor dialect" and "courtroom cross-examination." At k=100, clusters average 4,508 books each (of 9,766), confirming corpus-wide patterns. Direct comparison with embedding-similarity clustering shows that raw embeddings group by topic while association-space clusters group by function, register, and literary tradition. Unseen novels are assigned to existing clusters without retraining; the association model concentrates each novel into a selective subset of coherent clusters, while raw embedding assignment saturates nearly all clusters. Validation controls address positional, length, and book-concentration confounds. The method extends Predictive Associative Memory (PAM, arXiv:2602.11322) from episodic recall to concept formation: where PAM recalls specific associations, multi-epoch contrastive training under compression extracts structural patterns that transfer to unseen texts, the same framework producing qualitatively different behaviour in a different regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。