用范畴论解析文档结构,实现信息度量、摘要与模型自提升。
Document Understanding, Measurement, and Manipulation Using Category Theory
- 将文档建模为问答对范畴,数学化表达其内在结构。
- 通过正交化分离文档信息,实现非重叠内容分解与量化。
- 支持摘要优化与文档扩展,适合多模态理解与模型改进场景。
我们应用范畴论提取多模态文档结构,进而发展信息论度量、内容摘要与扩展方法,并实现大预训练模型的自监督优化。首先,将文档表示为问答对范畴;其次,提出正交化流程,将一个或多个文档的信息分解为互不重叠的部分。前两步所提取的结构使我们能够测量并枚举文档中的信息量。基于此,我们开发了新的摘要技术,并解决了一个新问题——诠释(exegesis),实现对原文档的合理扩展。问答对方法支持对摘要技术的新型率失真分析。我们使用大预训练模型实现这些技术,并提出整体数学框架的多模态扩展。最后,我们设计了一种新颖的自监督方法(基于RLVR),利用一致性约束(如可组合性、特定操作下的封闭性)来提升大预训练模型性能,这些约束源于范畴论框架本身。
原文摘要 · Abstract (English)
We apply category theory to extract multimodal document structure which leads us to develop information theoretic measures, content summarization and extension, and self-supervised improvement of large pretrained models. We first develop a mathematical representation of a document as a category of question-answer pairs. Second, we develop an orthogonalization procedure to divide the information contained in one or more documents into non-overlapping pieces. The structures extracted in the first and second steps lead us to develop methods to measure and enumerate the information contained in a document. We also build on those steps to develop new summarization techniques, as well as to develop a solution to a new problem viz. exegesis resulting in an extension of the original document. Our question-answer pair methodology enables a novel rate distortion analysis of summarization techniques. We implement our techniques using large pretrained models, and we propose a multimodal extension of our overall mathematical framework. Finally, we develop a novel self-supervised method using RLVR to improve large pretrained models using consistency constraints such as composability and closure under certain operations that stem naturally from our category theoretic framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。