arXiv:2508.19093cs.CL2025-08

用自然语言搜索艺术文物归属史,突破传统元数据限制。

Retrieval-Augmented Generation for Natural Language Art Provenance Searches in the Getty Provenance Index

  • 结合语义检索与生成摘要,支持多语言查询
  • 在1万条德国拍卖记录上验证,显著提升检索效率
  • 适合历史学家和文化遗产从业者做敏感研究

本研究提出一种检索增强生成(RAG)框架,用于艺术归属研究,聚焦于盖蒂归属索引(Getty Provenance Index)。归属研究旨在厘清艺术品的拥有历史,对验证真伪、支持返还与法律主张、理解文化历史背景至关重要。然而,档案数据分散且多语言,现有检索系统依赖精确元数据,难以支持探索性查询。本文方法通过语义检索与上下文摘要,实现自然语言和多语言搜索,降低对元数据结构的依赖。我们基于盖蒂归属索引-德国销售样本(10,000条记录)评估RAG在检索与总结拍卖记录方面的能力。结果表明,该方法为艺术市场档案提供可扩展的导航方案,是历史学家与文化遗产专业人员开展敏感研究的实用工具。

原文摘要 · Abstract (English)

This research presents a Retrieval-Augmented Generation (RAG) framework for art provenance studies, focusing on the Getty Provenance Index. Provenance research establishes the ownership history of artworks, which is essential for verifying authenticity, supporting restitution and legal claims, and understanding the cultural and historical context of art objects. The process is complicated by fragmented, multilingual archival data that hinders efficient retrieval. Current search portals require precise metadata, limiting exploratory searches. Our method enables natural-language and multilingual searches through semantic retrieval and contextual summarization, reducing dependence on metadata structures. We assess RAG's capability to retrieve and summarize auction records using a 10,000-record sample from the Getty Provenance Index - German Sales. The results show this approach provides a scalable solution for navigating art market archives, offering a practical tool for historians and cultural heritage professionals conducting historically sensitive research.

艺术溯源检索生成文化遗产

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。