arXiv:2602.19133cs.CL2026-02

构建艺术图像描述的细粒度命名实体与关系数据集,助力文物信息自动化提取。

A Dataset for Named Entity Recognition and Relation Extraction from Art-historical Image Descriptions

  • 三层次标注:元数据、内容主题、指代消解,覆盖37类实体
  • 支持跨模态知识图谱构建,可直接对接Wikidata
  • 适配大模型零样本/少样本训练,含真实图片与文献元数据

本文提出FRAME(细粒度艺术历史元数据与实体识别)数据集,用于命名实体识别(NER)与关系抽取(RE)。数据来自博物馆目录、拍卖清单、开放平台及学术数据库,经筛选确保每段文本聚焦单一艺术品,并明确描述其材质、构图或图像志。该数据集采用三层次站位标注:元数据层(对象属性)、内容层(描绘主体与母题)、共指层(重复提及关联),共标注37类实体,实体间通过类型化关系连接。实体类型与Wikidata对齐,支持命名实体链接(NEL)与下游知识图谱构建。数据以UIMA XMI CAS格式发布,附带图像和书目元数据,可用于基准测试与微调各类NER/RE系统,包括大语言模型的零样本与少样本设置。

原文摘要 · Abstract (English)

This paper introduces FRAME (Fine-grained Recognition of Art-historical Metadata and Entities), a manually annotated dataset of art-historical image descriptions for Named Entity Recognition (NER) and Relation Extraction (RE). Descriptions were collected from museum catalogs, auction listings, open-access platforms, and scholarly databases, then filtered to ensure that each text focuses on a single artwork and contains explicit statements about its material, composition, or iconography. FRAME provides stand-off annotations in three layers: a metadata layer for object-level properties, a content layer for depicted subjects and motifs, and a co-reference layer linking repeated mentions. Across layers, entity spans are labeled with 37 types and connected by typed RE links between mentions. Entity types are aligned with Wikidata to support Named Entity Linking (NEL) and downstream knowledge-graph construction. The dataset is released as UIMA XMI Common Analysis Structure (CAS) files with accompanying images and bibliographic metadata, and can be used to benchmark and fine-tune NER and RE systems, including zero- and few-shot setups with Large Language Models (LLMs).

命名实体艺术数据关系抽取知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。