利用画作裂纹特征实现历史油画多模态图像自动对齐,提升精度与效率。
Coarse-to-Fine Non-rigid Multi-modal Image Registration for Historical Panel Paintings based on Crack Structures
- 基于裂纹结构的稀疏关键点检测与图神经网络匹配,实现跨模态特征对应
- 提出分阶段精化方法,成功对齐不同分辨率图像,最高达1024×1024像素
- 构建首个含大量标注的关键点数据集,适合艺术分析与计算机视觉研究者
历史木板绘画的艺术技术研究依赖于多种成像方式获取的数据,包括可见光摄影、红外反射成像、紫外荧光摄影、X射线成像和宏观摄影。为实现全面分析,需对多模态图像进行像素级对齐,但目前仍主要依赖人工操作。多模态图像配准可显著减少人工工作量,提高速度与精度。由于图像分辨率差异大、尺寸巨大、存在非刚性形变且模态间内容各异,配准极具挑战。为此,我们提出一种基于稀疏关键点与薄板样条的粗到精非刚性配准方法。历史绘画表面的细密裂纹(craquelure)在所有成像系统中均可见,适合作为配准特征。在单阶段非刚性配准中,采用卷积神经网络联合检测与描述裂纹关键点,并通过基于块的图神经网络进行描述子匹配,结合局部同伦重投影误差过滤匹配。针对粗到精配准,引入新颖的多层级关键点精化策略,实现多分辨率图像的逐级高精度对齐。我们构建了一个包含大量关键点标注的多模态木板画数据集,测试集涵盖五种成像域及不同分辨率。消融实验证明各模块有效性。相比现有关键点与密集匹配方法及精化策略,本方法在配准性能上表现最佳。
原文摘要 · Abstract (English)
Art technological investigations of historical panel paintings rely on acquiring multi-modal image data, including visual light photography, infrared reflectography, ultraviolet fluorescence photography, x-radiography, and macro photography. For a comprehensive analysis, the multi-modal images require pixel-wise alignment, which is still often performed manually. Multi-modal image registration can reduce this laborious manual work, is substantially faster, and enables higher precision. Due to varying image resolutions, huge image sizes, non-rigid distortions, and modality-dependent image content, registration is challenging. Therefore, we propose a coarse-to-fine non-rigid multi-modal registration method efficiently relying on sparse keypoints and thin-plate-splines. Historical paintings exhibit a fine crack pattern, called craquelure, on the paint layer, which is captured by all image systems and is well-suited as a feature for registration. In our one-stage non-rigid registration approach, we employ a convolutional neural network for joint keypoint detection and description based on the craquelure and a graph neural network for descriptor matching in a patch-based manner, and filter matches based on homography reprojection errors in local areas. For coarse-to-fine registration, we introduce a novel multi-level keypoint refinement approach to register mixed-resolution images up to the highest resolution. We created a multi-modal dataset of panel paintings with a high number of keypoint annotations, and a large test set comprising five multi-modal domains and varying image resolutions. The ablation study demonstrates the effectiveness of all modules of our refinement method. Our proposed approaches achieve the best registration results compared to competing keypoint and dense matching methods and refinement methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。