arXiv:2507.04377cs.CVcs.CL2025-07被引 1

用视觉语言模型解析墓碑,让历史文字重获新生

Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions

  • 结合图像与文本的多模态框架,生成结构化语义表示
  • 准确率从36.1提升至89.5,显著优于传统OCR方法
  • 支持多种文化语言,可应对破损、模糊等现实问题

墓碑是蕴含个人生命历程、集体记忆、历史叙事与艺术表达的重要文化遗产。然而,许多墓碑正面临物理风化、人为破坏、环境退化及政治变迁等严峻保护挑战。本文提出一种新型多模态数字重建框架,旨在提升墓碑内容的解读、组织与检索能力。该方法利用视觉语言模型(VLMs)将墓碑图像转化为结构化的墓碑意义表示(TMRs),同时融合检索增强生成(RAG)技术,引入地名、职业代码与本体概念等外部知识。相比传统OCR流水线,该方法将解析准确率由F1分数36.1提升至89.5。我们还在多样语言与文化背景的墓碑上评估模型鲁棒性,并通过图像融合模拟物理损伤,测试其在噪声或损坏条件下的表现。本工作首次尝试以大视觉语言模型形式形式化墓碑理解,为文化遗产保护提供新路径。

原文摘要 · Abstract (English)

Tombstones are historically and culturally rich artifacts, encapsulating individual lives, community memory, historical narratives and artistic expression. Yet, many tombstones today face significant preservation challenges, including physical erosion, vandalism, environmental degradation, and political shifts. In this paper, we introduce a novel multi-modal framework for tombstones digitization, aiming to improve the interpretation, organization and retrieval of tombstone content. Our approach leverages vision-language models (VLMs) to translate tombstone images into structured Tombstone Meaning Representations (TMRs), capturing both image and text information. To further enrich semantic parsing, we incorporate retrieval-augmented generation (RAG) for integrate externally dependent elements such as toponyms, occupation codes, and ontological concepts. Compared to traditional OCR-based pipelines, our method improves parsing accuracy from an F1 score of 36.1 to 89.5. We additionally evaluate the model's robustness across diverse linguistic and cultural inscriptions, and simulate physical degradation through image fusion to assess performance under noisy or damaged conditions. Our work represents the first attempt to formalize tombstone understanding using large vision-language models, presenting implications for heritage preservation.

多模态遗产保护语义解析视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。