用少量数据就能生成红外、偏振等多模态新视角图像
Learning Spectral and Polarimetric Clues for One-to-Multimodal Novel View Synthesis

- 通过多模态预训练学习不同影像间的关联规律
- 仅用RGB图像微调,即可还原无传感器输入的红外/偏振图像
- 适合缺乏多模态传感器但需生成特殊成像的应用场景
神经渲染技术可精确重建3D场景的几何与颜色外观。尽管已有方法扩展至多光谱、红外或偏振等多模态数据,但这些方法均需昂贵传感器和校准设备来为每个新场景采集多模态帧。本文提出SPoILeR——一种新型隐式多模态表示方法,可在仅有RGB图像或极少量额外模态数据的情况下,生成一致的多视角非传统模态图像。模型通过多模态预训练学习各模态间关联,使在仅以RGB图像监督的微调阶段,即可准确预测红外、偏振及多光谱图像。实验表明,该方法能在无任何对应传感器输入的前提下,精准生成上述多种模态的新视角图像。
原文摘要 · Abstract (English)
Neural rendering techniques allow for accurate reconstruction of the geometry and color appearance of 3D scenes. Some methods have extended their use to additional imaging modalities, such as multispectral, infrared, or polarimetric data. However, all of these approaches require expensive sensors and calibrated setups to capture new multimodal frames for each new scene. We propose Spectral and Polarimetric Implicit Learned Representation (SPoILeR), a novel method to obtain multi-view consistent renderings of unconventional modalities for scenes where either only RGB frames or very few of the additional modalities are available. Thanks to a multimodal pre-training phase, the model learns the mutual correlation between different modalities. This step allows predicting accurate renderings of unconventional modalities during a fine-tuning phase supervised only by RGB images. Experimental results show that the approach can accurately render infrared, polarimetric, and multispectral frames for scenes where no input sample captured by these types of sensors is provided.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。