用时空变压器提升低光下的3D重建质量,适用于可见与非可见光成像。
3D Reconstruction from Transient Measurements with Time-Resolved Transformer
- 设计双注意力机制的时空变压器,兼顾局部与全局特征关联。
- 在真实与合成数据上均超越现有方法,尤其在高噪声长距场景下表现优异。
- 开源代码与大规模合成/实测数据集,助力后续研究。
瞬态测量由时间分辨系统捕获,广泛应用于光子效率高的三维重建任务,包括视线内(LOS)和非视线内(NLOS)成像。然而,由于传感器量子效率低、噪声水平高,尤其在远距离或复杂场景下,三维重建仍面临挑战。为提升光子高效成像中的三维重建性能,我们提出一种通用的时间分辨变压器(TRT)架构。不同于面向高维数据的现有变压器,TRT采用两种专为时空瞬态测量设计的注意力机制:时空自注意力编码器通过分块或降采样输入特征,探索瞬态数据中的局部与全局相关性;时空交叉注意力解码器则在标记空间中融合局部与全局特征,生成具有强表达能力的深层特征。基于TRT,我们构建了两个特定任务版本:用于视线成像的TRT-LOS和用于非视线成像的TRT-NLOS。大量实验表明,二者在合成数据和不同成像系统采集的真实数据上均显著优于现有方法。此外,我们还构建了一个大规模、高分辨率的合成LOS数据集,涵盖多种噪声水平,并使用自研成像系统采集了一组真实世界非视线测量数据,提升了该领域的数据多样性。代码与数据集已公开于 https://github.com/Depth2World/TRT。
原文摘要 · Abstract (English)
Transient measurements, captured by the timeresolved systems, are widely employed in photon-efficient reconstruction tasks, including line-of-sight (LOS) and non-line-of-sight (NLOS) imaging. However, challenges persist in their 3D reconstruction due to the low quantum efficiency of sensors and the high noise levels, particularly for long-range or complex scenes. To boost the 3D reconstruction performance in photon-efficient imaging, we propose a generic Time-Resolved Transformer (TRT) architecture. Different from existing transformers designed for high-dimensional data, TRT has two elaborate attention designs tailored for the spatio-temporal transient measurements. Specifically, the spatio-temporal self-attention encoders explore both local and global correlations within transient data by splitting or downsampling input features into different scales. Then, the spatio-temporal cross attention decoders integrate the local and global features in the token space, resulting in deep features with high representation capabilities. Building on TRT, we develop two task-specific embodiments: TRT-LOS for LOS imaging and TRT-NLOS for NLOS imaging. Extensive experiments demonstrate that both embodiments significantly outperform existing methods on synthetic data and real-world data captured by different imaging systems. In addition, we contribute a large-scale, high-resolution synthetic LOS dataset with various noise levels and capture a set of real-world NLOS measurements using a custom-built imaging system, enhancing the data diversity in this field. Code and datasets are available at https://github.com/Depth2World/TRT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。