arXiv:2512.07452cs.IR2025-12

用大模型把戏单变结构化数据,让戏剧遗产可查可用

From Show Programmes to Data: Designing a Workflow to Make Performing Arts Ephemera Accessible Through Language Models

  • 结合多模态大模型与知识图谱,自动解析戏单内容
  • 戏单信息提取准确率超98%,实现高质量数据转化
  • 适合文化遗产机构和数字人文研究者使用

众多文史机构收藏了大量戏剧节目单,但因版式复杂、缺乏结构化元数据而难以利用。本文提出一套工作流,结合多模态大语言模型、基于本体的推理模型及定制化的链接艺术框架,将数字化和原生数字戏单转化为结构化数据。视觉-语言模型在戏单内容解析与转录中表现优异,准确率超过98%。为解决语义标注难题,我们采用强化学习训练了一个推理模型POntAvignon,同时引入形式与语义奖励机制,实现自动化RDF三元组生成,并支持与现有知识图谱对齐。基于阿维尼翁节档案的案例研究验证了该方法在大规模、本体驱动的表演艺术数据分析中的潜力。结果为可互操作、可解释、可持续的计算戏剧史学开辟新路径。

原文摘要 · Abstract (English)

Many heritage institutions hold extensive collections of theatre programmes, which remain largely underused due to their complex layouts and lack of structured metadata. In this paper, we present a workflow for transforming such documents into structured data using a combination of multimodal large language models (LLMs), an ontology-based reasoning model, and a custom extension of the Linked Art framework. We show how vision-language models can accurately parse and transcribe born-digital and digitised programmes, achieving over 98% of correct extraction. To overcome the challenges of semantic annotation, we train a reasoning model (POntAvignon) using reinforcement learning with both formal and semantic rewards. This approach enables automated RDF triple generation and supports alignment with existing knowledge graphs. Through a case study based on the Festival d'Avignon corpus, we demonstrate the potential for large-scale, ontology-driven analysis of performing arts data. Our results open new possibilities for interoperable, explainable, and sustainable computational theatre historiography.

文化遗产大模型知识图谱戏剧史

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。