用信息论方法识别文本中重复性强的结构片段,适合研究多作者文献。
An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data
- 基于微分自信息构建连续化信息度量,捕捉文本中的规律性模式。
- 在希伯来圣经上成功划分出多个风格层,揭示潜在作者分界。
- 无需标签即可分析稀疏高维文本,适用于神经嵌入等现代表示。
文本,无论是文学还是历史性的,其结构与风格都受到目的、作者和文化背景的影响。公式化文本以重复和表达受限为特征,其信息内容(按香农定义)通常不同于更具动态性的作品。在历史文献中,尤其是多作者文本如《希伯来圣经》,识别此类模式有助于理解其起源、目的与传播过程。本研究旨在通过分析重复短语、句法结构和风格标记,识别公式化聚类——即表现出系统重复与结构约束的段落。然而,在无监督条件下区分公式化与非公式化成分在高维、样本稀疏的数据集中构成计算挑战。为此,我们提出一种信息论算法,利用加权自信息分布恢复文本中的结构分区。所得聚类通过自信息特征与典型重复特征进行解释。通过将经典离散自信息扩展为基于多元高斯分布的连续形式,该方法可适用于多种文本表示,包括在高斯先验下的神经嵌入。应用于《希伯来圣经》的假定作者划分,该方法成功分离出风格层次,并提供了一种量化文本分层的框架。该方法增强了对复杂作者与编辑过程中形成的文本组合模式的分析能力,深化了对文本文学与文化演变的理解。
原文摘要 · Abstract (English)
Texts, whether literary or historical, exhibit structural and stylistic patterns shaped by their purpose, authorship, and cultural context. Formulaic texts, which are characterized by repetition and constrained expression, tend to differ in their \textit{information content} (as defined by Shannon) compared to more dynamic compositions. Identifying such patterns in historical documents, particularly multi-author texts like the Hebrew Bible, provides insights into their origins, purpose, and transmission. This study aims to identify formulaic clusters: sections exhibiting systematic repetition and structural constraints, by analyzing recurring phrases, syntactic structures, and stylistic markers. However, distinguishing formulaic from non-formulaic elements in an unsupervised manner poses a computational challenge, especially in high-dimensional, sample-poor data sets where patterns must be inferred without predefined labels. To address this, we develop an information-theoretic algorithm that uses weighted \textit{self-information} distributions to recover structured partitions in text. The resulting clusters are interpreted from their self-information profiles and characteristic recurring features. By extending classical discrete self-information measures to a continuous formulation based on differential self-information in multivariate Gaussian distributions, our method remains applicable across various textual representations, including neural embeddings under Gaussian priors. Applied to hypothesized authorial divisions in the Hebrew Bible, our approach isolates stylistic layers and provides a quantitative framework for textual stratification. This method enhances our ability to analyze compositional patterns, offering deeper insights into the literary and cultural evolution of texts shaped by complex authorship and editorial processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。