用可学习关键点和扩散模型重建破损壁画,精度远超现有方法。
ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco Reconstruction
- 将碎片表示为轮廓关键点,用图网络筛选重要点降低复杂度。
- 在真实破损壁画上实现旋转误差下降57%、平移误差下降87%。
- 适合文物修复、复杂几何重构等现实场景,支持多模态数据融合。
图像重装任务在考古、基因组学和分子对接等领域具有重要意义,需精确对齐并定位碎片以恢复原始结构。本文针对当前深度学习方法在可扩展性、多模态融合及真实场景适用性方面的局限,提出ReassembleNet:通过将每个输入碎片表示为一组轮廓关键点,并利用受图神经网络池化启发的机制学习选择最具信息量的关键点,有效降低计算复杂度,同时整合几何与纹理等多模态特征。模型在半合成数据集上预训练后,采用基于扩散的位姿估计恢复原始结构。实验表明,在旋转误差(RMSE)和翻译误差上分别优于先前方法57%和87%。
原文摘要 · Abstract (English)
The task of reassembly is a significant challenge across multiple domains, including archaeology, genomics, and molecular docking, requiring the precise placement and orientation of elements to reconstruct an original structure. In this work, we address key limitations in state-of-the-art Deep Learning methods for reassembly, namely i) scalability; ii) multimodality; and iii) real-world applicability: beyond square or simple geometric shapes, realistic and complex erosion, or other real-world problems. We propose ReassembleNet, a method that reduces complexity by representing each input piece as a set of contour keypoints and learning to select the most informative ones by Graph Neural Networks pooling inspired techniques. ReassembleNet effectively lowers computational complexity while enabling the integration of features from multiple modalities, including both geometric and texture data. Further enhanced through pretraining on a semi-synthetic dataset. We then apply diffusion-based pose estimation to recover the original structure. We improve on prior methods by 57% and 87% for RMSE Rotation and Translation, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。