用传感器与图像融合预测古迹退化,准确率提升43%。
A Multimodal Approach to Heritage Preservation in the Context of Climate Change
- 轻量级多模态架构融合温湿度与图像数据
- 在37个样本上达76.9%准确率,优于现有模型
- 适合数据稀缺场景的遗产保护智能系统
文化遗产地因气候变化正加速退化,传统监测依赖单一模态(仅视觉或环境传感器),难以捕捉环境压力与材料劣化间的复杂关系。本文提出一种轻量级多模态架构,融合传感器数据(温度、湿度)与视觉图像,以预测古迹退化程度。模型基于PerceiverIO,采用两项创新:(1) 简化编码器(64维潜在空间),防止小样本(n=37)过拟合;(2) 自适应Barlow Twins损失,促进模态互补而非冗余。在斯特拉斯堡大教堂数据上,模型准确率达76.9%,比标准多模态架构(VisualBERT、Transformer)高43%,比原版PerceiverIO高25%。消融实验显示,仅传感器为61.5%,仅图像为46.2%,验证了多模态协同效应。系统性超参数研究发现,适度相关性目标τ=0.3时最优,准确率达69.2%(τ=0.1/0.5/0.7:53.8%;τ=0.9:61.5%)。该工作表明,结构简化结合对比正则化可在数据稀缺的遗产监测中实现有效多模态学习,为人工智能驱动的保护决策系统提供基础。
原文摘要 · Abstract (English)
Cultural heritage sites face accelerating degradation due to climate change, yet tradi- tional monitoring relies on unimodal analysis (visual inspection or environmental sen- sors alone) that fails to capture the complex interplay between environmental stres- sors and material deterioration. We propose a lightweight multimodal architecture that fuses sensor data (temperature, humidity) with visual imagery to predict degradation severity at heritage sites. Our approach adapts PerceiverIO with two key innovations: (1) simplified encoders (64D latent space) that prevent overfitting on small datasets (n=37 training samples), and (2) Adaptive Barlow Twins loss that encourages modality complementarity rather than redundancy. On data from Strasbourg Cathedral, our model achieves 76.9% accu- racy, a 43% improvement over standard multimodal architectures (VisualBERT, Trans- former) and 25% over vanilla PerceiverIO. Ablation studies reveal that sensor-only achieves 61.5% while image-only reaches 46.2%, confirming successful multimodal synergy. A systematic hyperparameter study identifies an optimal moderate correlation target (τ =0.3) that balances align- ment and complementarity, achieving 69.2% accuracy compared to other τ values (τ =0.1/0.5/0.7: 53.8%, τ =0.9: 61.5%). This work demonstrates that architectural sim- plicity combined with contrastive regularization enables effective multimodal learning in data-scarce heritage monitoring contexts, providing a foundation for AI-driven con- servation decision support systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。