arXiv:2602.20417cs.CV2026-02

用大模型复原极低光下的光子级图像,让弱光成像更清晰。

gQIR: Generative Quanta Image Reconstruction

  • 将扩散模型适配到光子稀疏场景,融合时空推理与噪声建模。
  • 在真实色度SPAD数据集上,重建质量显著优于传统和现有方法。
  • 适合做超灵敏成像、低光照视觉的科研与工程人员参考。

从极少光子中获取高质量图像是计算成像的核心挑战。单光子雪崩二极管(SPAD)传感器可在传统相机失效的极端条件下实现高质成像,但原始的‘量子帧’仅包含稀疏、噪声严重的二值光子检测。从一系列此类帧中恢复出连贯图像,需同时处理对齐、去噪和解马赛克(彩色)问题,而其噪声统计远超标准修复流程或现代生成模型的假设范围。本文提出一种方法,将大规模文本到图像隐空间扩散模型适配至量子帧成像的光子受限领域。该方法利用互联网规模扩散模型的结构与语义先验,引入机制以应对伯努利光子统计特性。通过融合隐空间修复与帧间时空推理,重建结果在光度保真度和感知质量上均表现优异,即使在高速运动下亦然。我们在合成基准和新建立的真实世界数据集上评估该方法,包括首个彩色SPAD帧数据集及具有挑战性的‘形变(XD)’视频基准。在所有测试场景中,本方法在感知质量上显著超越经典与现代学习基线,证明了将大规模生成先验迁移至极端光子受限传感的潜力。代码见:https://github.com/Aryan-Garg/gQIR。

原文摘要 · Abstract (English)

Capturing high-quality images from only a few detected photons is a fundamental challenge in computational imaging. Single-photon avalanche diode (SPAD) sensors promise high-quality imaging in regimes where conventional cameras fail, but raw \emph{quanta frames} contain only sparse, noisy, binary photon detections. Recovering a coherent image from a burst of such frames requires handling alignment, denoising, and demosaicing (for color) under noise statistics far outside those assumed by standard restoration pipelines or modern generative models. We present an approach that adapts large text-to-image latent diffusion models to the photon-limited domain of quanta burst imaging. Our method leverages the structural and semantic priors of internet-scale diffusion models while introducing mechanisms to handle Bernoulli photon statistics. By integrating latent-space restoration with burst-level spatio-temporal reasoning, our approach produces reconstructions that are both photometrically faithful and perceptually pleasing, even under high-speed motion. We evaluate the method on synthetic benchmarks and new real-world datasets, including the first color SPAD burst dataset and a challenging \textit{Deforming (XD)} video benchmark. Across all settings, the approach substantially improves perceptual quality over classical and modern learning-based baselines, demonstrating the promise of adapting large generative priors to extreme photon-limited sensing. Code at \href{https://github.com/Aryan-Garg/gQIR}{https://github.com/Aryan-Garg/gQIR}.

低光成像扩散模型SPAD图像重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。