arXiv:2607.18907cs.CV2026-07

用合成图像提升艺术作品识别准确率,解决真实拍摄场景差异问题。

SynGallery: A Synthetic Gallery of Real Paintings for Instance-Level Artwork Recognition

论文配图:SynGallery: A Synthetic Gallery of Real Paintings for Instance-Level Artwork Recognition
图 1 · 摘自论文原文
  • 从真实画作目录图出发,生成带复杂视角和光照变化的合成画廊视图。
  • 仅用合成数据训练可使识别准确率从67.18提升至73.47 GAP⁻。
  • 适用于需要高鲁棒性图像匹配的博物馆智能导览与艺术检索系统。

实例级艺术作品识别需将游客手持拍摄的照片与大型博物馆藏品中的特定作品匹配。该任务困难在于,训练数据通常为清晰的目录图像,而测试查询则存在斜视角、灯光、反光、画框等场景级差异。我们提出SynGallery,一个无需额外采集真实照片的合成画廊数据集。基于真实画作的目录图像,将每幅作品置于程序生成的3D画廊场景中,从多个视角在不同几何与外观条件下渲染,同时保留原作的精确身份。该数据集包含来自Met基准的4,898幅画作共24,490个渲染视图。实验表明,这些合成视图提供的训练信号强于对应的棚拍照片。在相同训练样本数下,仅使用SynGallery训练可使艺术作品识别准确率从67.18提升至73.47 GAP⁻。当加入完整Met训练集时,基准协议性能从35.97提升至38.48 GAP。消融实验显示,增益主要来自场景级视角变化而非摄影写实性:将五种渲染视角缩减为单一正面视图会大幅降低性能;而模拟模糊、传感器噪声和压缩伪影则持续损害表现。

原文摘要 · Abstract (English)

Instance-level artwork recognition requires matching a handheld visitor photograph to a specific work in a large museum collection. This is challenging because painting datasets typically provide clean catalog images for training, while test queries are captured under oblique viewpoints, gallery lighting, reflections, frames, and other scene-level variations. We present SynGallery, a synthetic gallery dataset for artwork retrieval that addresses this gap without collecting additional real photographs. Starting from catalog images of real paintings, we place each artwork into a procedurally generated 3D gallery scene and render it from multiple viewpoints under varied geometric and appearance conditions, while preserving the exact identity of the original work. The resulting dataset contains 24,490 rendered views of 4,898 paintings from the Met benchmark. We show that these synthetic views provide a stronger training signal than the corresponding studio photographs. At the same number of training data points, training only on SynGallery improves art painting recognition from 67.18 to 73.47 GAP$^-$. When added to the full Met training set, SynGallery improves the published benchmark protocol from 35.97 to 38.48 GAP. Ablation experiments show that the gain comes from scene-level view variation rather than photographic realism: reducing the five rendered viewpoints to a single frontal view removes most of the improvement, while simulating capture artifacts such as blur, sensor noise, and image compression consistently reduces performance.

艺术识别合成数据图像检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。