arXiv:2503.08710eess.IVcs.CV2025-03被引 12

用14亿参数大模型提升鬼成像重建质量,抗噪强、远距离成像效果好。

Large model enhanced computational ghost imaging

  • 构建1.4亿参数大模型GILM,融合鬼成像物理机制与多头注意力。
  • 在52米水下实现清晰成像,信号波动分析能力显著提升重建精度。
  • 适合低光照、散射介质等复杂场景的高鲁棒性图像恢复应用。

鬼成像(GI)通过一维桶信号与二维光场信息的高阶相关性实现二维图像重建,尤其在散射介质中高效收集光子时表现出增强的检测灵敏度和高质量成像能力。近年来研究表明,深度学习(DL)可显著提升鬼成像重建质量。随着SDXL、GPT-4等大模型的出现,传统DL在参数量与架构上的限制被突破,使模型能全面捕捉特征序列中各位置间的复杂关系。本研究首次提出具备14亿参数的大型成像模型GILM,融合鬼成像物理原理。GILM采用跳跃连接缓解深层网络梯度爆炸问题,确保充分捕捉单像素测量间的复杂相关性;同时利用多头注意力机制学习像素点间空间依赖关系,促进完整物体信息提取。实验验证包括模拟物成像、自由空间成像及水下52米处目标成像,结果表明GILM能有效分析采集信号的波动趋势,优化从原始数据中恢复物体图像的过程。

原文摘要 · Abstract (English)

Ghost imaging (GI) achieves 2D image reconstruction through high-order correlation of 1D bucket signals and 2D light field information, particularly demonstrating enhanced detection sensitivity and high-quality image reconstruction via efficient photon collection in scattering media. Recent investigations have established that deep learning (DL) can substantially enhance the ghost imaging reconstruction quality. Furthermore, with the emergence of large models like SDXL, GPT-4, etc., the constraints of conventional DL in parameters and architecture have been transcended, enabling models to comprehensively explore relationships among all distinct positions within feature sequences. This paradigm shift has significantly advanced the capability of DL in restoring severely degraded and low-resolution imagery, making it particularly advantageous for noise-robust image reconstruction in GI applications. In this paper, we propose the first large imaging model with 1.4 billion parameters that incorporates the physical principles of GI (GILM). The proposed GILM implements a skip connection mechanism to mitigate gradient explosion challenges inherent in deep architectures, ensuring sufficient parametric capacity to capture intricate correlations among object single-pixel measurements. Moreover, GILM leverages multi-head attention mechanism to learn spatial dependencies across pixel points during image reconstruction, facilitating the extraction of comprehensive object information for subsequent reconstruction. We validated the effectiveness of GILM through a series of experiments, including simulated object imaging, imaging objects in free space, and imaging object located 52 meters away in underwater environment. The experimental results show that GILM effectively analyzes the fluctuation trends of the collected signals, thereby optimizing the recovery of the object's image from the acquired data.

鬼成像大模型水下成像注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。