用视觉语言模型解析皮尔士手稿中的图文混合内容
Moving Pictures of Thought: Extracting Visual Knowledge in Charles S. Peirce's Manuscripts with Vision-Language Models
- 分块手稿页面并关联标注,输入视觉模型分析图示
- 基于符号学框架提取图示核心知识并生成简明描述
- 将结果构建成知识图谱,实现图文内容结构化
图表在众多学科中至关重要却未被充分研究,其图像形式给视觉研究、跨媒介分析及文本数字化流程带来挑战。查尔斯·皮尔士始终倡导以图表作为推理与解释的关键工具。其手稿常融合文字与复杂视觉元素,构成异质材料研究的难题。本初步研究探讨视觉语言模型(VLMs)是否能有效识别并解读此类混合页面。首先,提出工作流:(i) 分割手稿页面布局,(ii) 将各片段与符合IIIF标准的注释重新关联,(iii) 将含图段落提交至VLM。同时,基于皮尔士符号学框架设计提示词,提取图示关键知识并生成简洁标题。最终,将这些标题整合进知识图谱,实现复合文献中图示内容的结构化表示。
原文摘要 · Abstract (English)
Diagrams are crucial yet underexplored tools in many disciplines, demonstrating the close connection between visual representation and scholarly reasoning. However, their iconic form poses obstacles to visual studies, intermedial analysis, and text-based digital workflows. In particular, Charles S. Peirce consistently advocated the use of diagrams as essential for reasoning and explanation. His manuscripts, often combining textual content with complex visual artifacts, provide a challenging case for studying documents involving heterogeneous materials. In this preliminary study, we investigate whether Visual Language Models (VLMs) can effectively help us identify and interpret such hybrid pages in context. First, we propose a workflow that (i) segments manuscript page layouts, (ii) reconnects each segment to IIIF-compliant annotations, and (iii) submits fragments containing diagrams to a VLM. In addition, by adopting Peirce's semiotic framework, we designed prompts to extract key knowledge about diagrams and produce concise captions. Finally, we integrated these captions into knowledge graphs, enabling structured representations of diagrammatic content within composite sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。