arXiv:2605.23531cs.CV2026-05

用大模型引导像素级低光图像增强,去噪同时恢复细节

PixIE: Prompted Pixel-Space Low-Light Image Enhancement

论文配图:PixIE: Prompted Pixel-Space Low-Light Image Enhancement
图 1 · 摘自论文原文
  • 通过跨尺度去噪和提示调制块,融合大模型语义信息
  • 在LOLv2-Real上达到最佳PSNR、SSIM与LPIPS指标
  • 适合需要高保真与自然纹理的低光图像修复场景

低光图像普遍存在严重噪声、对比度丧失和语义模糊问题,增强需兼顾去噪与细节恢复。我们提出PixIE,一种前馈式像素空间低光图像增强框架,由基础模型(FM)进行语义提示。PixIE首先进行跨尺度去噪以抑制噪声并保留结构,随后利用提示像素块(PPBs)通过新颖的空间连续调制(SCMo)注入中间FM特征以细化细节。为提升多尺度像素空间注意力效率,引入空间-通道压缩(SCC),联合减少空间令牌网格与通道维度。进一步提出多感受野像素嵌入(MRPE),在语义提示前提供邻域感知的像素表示,增强对信号依赖性噪声的鲁棒性。在标准低光图像增强基准上实验表明,PixIE在挑战性的LOLv2-Real基准上取得最优性能,达到最高PSNR、SSIM与最低LPIPS。定性对比显示细节更锐利,纹理更自然一致,显著提升重建保真度与感知质量。

原文摘要 · Abstract (English)

Low-light images suffer from severe noise, contrast loss, and semantic ambiguity, making enhancement a joint problem of denoising and detail recovery. We propose PixIE, a feed-forward pixel-space LLIE framework semantically prompted by a foundation model (FM). PixIE first performs cross-scale denoising to suppress noise while preserving structure, then refines details using Prompted Pixel Blocks (PPBs), which inject intermediate FM features through a novel spatially continuous modulation (SCMo). To make pixel-space attention efficient across scales, we introduce Spatial-Channel Compaction (SCC), which jointly reduces the spatial token grid and channel dimension. We further propose Multi-Receptive-Field Pixel Embedding (MRPE) to provide neighborhood-aware pixel representations before semantic prompting, improving robustness to signal-dependent noise beyond point-wise embeddings. Experiments on standard LLIE benchmarks demonstrate state-of-the-art performance, achieving the best PSNR, SSIM, and LPIPS on the challenging LOLv2-Real benchmark. Qualitative comparisons further show sharper details with more natural and consistent textures, improving both reconstruction fidelity and perceptual quality.

低光增强像素空间大模型提示去噪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。