用自监督模型MARMOT提升非视域成像的重建效果
MARMOT: Masked Autoencoder for Modeling Transient Imaging
- 基于Transformer的掩码自编码器,通过扫描模式掩码学习部分遮挡的瞬态数据
- 在50万3D模型合成数据集上预训练,下游任务直接迁移特征或微调解码器
- 适用于非视域成像,可高效重建隐藏物体,适合计算机视觉与光子成像研究者
预训练模型在语言和视觉等领域取得显著成功。近期工作将预训练范式引入成像研究。瞬态成像是一种新型模态,通过高精度时间分辨传感器捕获物体的光子计数随到达时间的变化。特别是在非视域(NLOS)场景中,传感器可测量隐藏物体的瞬态信号。以往多数方法通过优化体密度或表面来重建隐藏物体,但未利用从数据集中学习到的先验知识。本文提出一种用于瞬态成像建模的掩码自编码器(MARMOT),以支持NLOS应用。MARMOT是基于Transformer的自监督模型,在大规模多样化的NLOS瞬态数据集上预训练。其采用扫描模式掩码(SPM),使未掩码部分等价于任意采样,从而预测完整测量结果。在包含50万3D模型的合成数据集TransVerse上预训练后,MARMOT可通过直接特征迁移或解码器微调适配下游成像任务。大量实验表明,MARMOT在定量与定性指标上均优于现有最先进方法。
原文摘要 · Abstract (English)
Pretrained models have demonstrated impressive success in many modalities such as language and vision. Recent works facilitate the pretraining paradigm in imaging research. Transients are a novel modality, which are captured for an object as photon counts versus arrival times using a precisely time-resolved sensor. In particular for non-line-of-sight (NLOS) scenarios, transients of hidden objects are measured beyond the sensor's direct line of sight. Using NLOS transients, the majority of previous works optimize volume density or surfaces to reconstruct the hidden objects and do not transfer priors learned from datasets. In this work, we present a masked autoencoder for modeling transient imaging, or MARMOT, to facilitate NLOS applications. Our MARMOT is a self-supervised model pretrianed on massive and diverse NLOS transient datasets. Using a Transformer-based encoder-decoder, MARMOT learns features from partially masked transients via a scanning pattern mask (SPM), where the unmasked subset is functionally equivalent to arbitrary sampling, and predicts full measurements. Pretrained on TransVerse-a synthesized transient dataset of 500K 3D models-MARMOT adapts to downstream imaging tasks using direct feature transfer or decoder finetuning. Comprehensive experiments are carried out in comparisons with state-of-the-art methods. Quantitative and qualitative results demonstrate the efficiency of our MARMOT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。