arXiv:2606.07753cs.CL2026-06

用大模型对文档集做结构化阅读,自动提取上千条观点并生成主题图谱。

ReadingMachine: A Computational Methodology for Structured Corpus Reading and Large-Scale Synthesis

  • 分阶段处理:先提取洞察、聚类语义、生成主题,再检测遗漏,全程可追踪。
  • 在152份政策文件中提取超1.75万条见解,生成结构化主题地图。
  • 适合需要深度分析海量文本的学者或政策研究者使用。

ReadingMachine 是一种用于结构化文献阅读的计算方法,利用大语言模型对整个文档集合执行受控阅读操作。该方法不依赖检索或递归摘要,而是将分析分解为可检查的阶段,包括洞察提取、语义聚类、主题生成和迭代遗漏检测。通过延迟不可逆压缩,并显式跟踪中间表示,该方法优先保证覆盖范围、可追溯性以及对大型语料库中分歧意见的保留。系统在包含152份工业政策文档的异构语料库上进行了演示,生成了超过17,500条提取的见解和一个结构化的主题地图。ReadingMachine 作为开源实验框架发布,支持大规模定性合成与语料分析。

原文摘要 · Abstract (English)

ReadingMachine is a computational methodology for structured corpus reading that uses large language models to perform bounded reading operations over entire document collections. Rather than relying on retrieval or recursive summarization, the approach decomposes analysis into inspectable stages including insight extraction, semantic clustering, theme generation, and iterative omission detection. By delaying irreversible compression and explicitly tracking intermediate representations, the method prioritizes coverage, traceability, and preservation of disagreement across large corpora. The system is demonstrated on a heterogeneous corpus of 152 industrial policy documents, producing more than 17,500 extracted insights and a structured thematic map. ReadingMachine is released as an open-source experimental framework for large-scale qualitative synthesis and corpus analysis.

文本分析大模型主题建模开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。