首个针对模拟媒介损伤检测的基准数据集,助力文化遗产保护。
ARTeFACT: Benchmarking Segmentation Models on Diverse Analogue Media Damage
- 构建包含15类损伤的1.1万+标注数据集,覆盖多种媒介与历史背景。
- 多模型测试显示:现有模型在跨媒介泛化上表现不佳。
- 提供文本提示与损伤描述,支持零样本与文本引导检测。
准确检测与分类绘画、照片、纺织品、马赛克和壁画等模拟媒介中的损伤,对文化遗产保护至关重要。尽管机器学习模型在已知退化过程时能有效修复,但即便经过监督训练,仍难以可靠预测损伤位置,因此损伤检测仍是难题。为此,我们提出ARTeFACT,首个针对多样模拟媒介损伤检测的数据集,包含超过11,000个标注,涵盖15类损伤,覆盖多种主题、媒介及历史来源。我们还提供了经人工验证的图像语义文本提示,并衍生出损伤的附加文本描述。我们在零样本、监督、无监督及文本引导设置下评估了CNN、Transformer、基于扩散的分割模型和基础视觉模型,揭示其在跨媒介泛化上的局限性。数据集已发布于https://daniela997.github.io/ARTeFACT/,是首个此类基准。
原文摘要 · Abstract (English)
Accurately detecting and classifying damage in analogue media such as paintings, photographs, textiles, mosaics, and frescoes is essential for cultural heritage preservation. While machine learning models excel in correcting degradation if the damage operator is known a priori, we show that they fail to robustly predict where the damage is even after supervised training; thus, reliable damage detection remains a challenge. Motivated by this, we introduce ARTeFACT, a dataset for damage detection in diverse types analogue media, with over 11,000 annotations covering 15 kinds of damage across various subjects, media, and historical provenance. Furthermore, we contribute human-verified text prompts describing the semantic contents of the images, and derive additional textual descriptions of the annotated damage. We evaluate CNN, Transformer, diffusion-based segmentation models, and foundation vision models in zero-shot, supervised, unsupervised and text-guided settings, revealing their limitations in generalising across media types. Our dataset is available at $\href{https://daniela997.github.io/ARTeFACT/}{https://daniela997.github.io/ARTeFACT/}$ as the first-of-its-kind benchmark for analogue media damage detection and restoration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。