arXiv:2603.11147cs.MMcs.CV2026-03

用本地部署模型自动为博物馆视频生成可搜索的元数据。

Catalogue Grounded Multimodal Attribution for Museum Video under Resource and Regulatory Constraints

  • 基于现有藏品库,用多轮流程提取视频中的艺术品并生成目录式描述。
  • 在绘画藏品上测试显示,能提升音视频档案可发现性且符合数据合规要求。
  • 适合需要数据主权和法规合规的博物馆、美术馆等高风险领域使用。

博物馆与画廊的音视频档案快速增长,但因缺乏一致且可检索的元数据,大量资料仍处于无法访问状态。现有归档方法依赖大量人工工作。本文通过自动化最耗时的环节——基于已有藏品数据库的馆藏风格元数据整理,提出面向博物馆音视频内容的“目录锚定多模态归属”方法。具体而言,采用开源、可本地部署的视频语言模型,设计多阶段流程:(i) 摘要视频中的艺术品;(ii) 生成目录风格描述与类别标签;(iii) 通过保守相似度匹配尝试归属作品标题与作者。早期在绘画藏品上的部署表明,该框架可在满足资源限制、数据主权及新兴监管要求的前提下,显著提升音视频档案的可发现性,为其他高风险领域中的应用驱动型机器学习提供可复用模板。

原文摘要 · Abstract (English)

Audiovisual (AV) archives in museums and galleries are growing rapidly, but much of this material remains effectively locked away because it lacks consistent, searchable metadata. Existing method for archiving requires extensive manual effort. We address this by automating the most labour intensive part of the workflow: catalogue style metadata curation for in gallery video, grounded in an existing collection database. Concretely, we propose catalogue-grounded multimodal attribution for museum AV content using an open, locally deployable video language model. We design a multi pass pipeline that (i) summarises artworks in a video, (ii) generates catalogue style descriptions and genre labels, and (iii) attempts to attribute title and artist via conservative similarity matching to the structured catalogue. Early deployments on a painting catalogue suggest that this framework can improve AV archive discoverability while respecting resource constraints, data sovereignty, and emerging regulation, offering a transferable template for application-driven machine learning in other high-stakes domains.

元数据生成多模态数据合规博物馆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。