用克利福德代数实现4K/8K低光图像实时增强,边缘设备毫秒级推理。
UHD Low-Light Image Enhancement via Real-Time Enhancement Methods with Clifford Information Fusion
- 基于克利福德代数的几何特征融合,提升高低频信息整合能力。
- 在4K/8K图像上实现毫秒级推理,优于现有最先进模型。
- 适合移动端或嵌入式设备部署,兼顾速度与图像保真度。
考虑到效率,超高清(UHD)低光图像恢复极具挑战性。现有基于Transformer架构或高维复数卷积神经网络的方法常受‘内存墙’瓶颈制约,难以在边缘设备上实现毫秒级推理。为此,我们提出一种基于二维欧几里得空间克利福德代数的实时超高清低光增强网络。首先构建四层渐进分辨率特征金字塔,通过高斯模糊核将输入图像分解为低频与高频结构成分,并采用基于深度可分离卷积的轻量级U-Net进行双分支特征提取。其次,为解决传统高低频特征融合导致的结构信息丢失和伪影问题,引入空间感知克利福德代数,将特征张量映射至多向量空间(标量、向量、二向量),利用克利福德相似性聚合特征,抑制噪声并保留纹理。重建阶段输出自适应伽马与增益图,通过受限非线性亮度调整实现物理合理的增强。结合FP16混合精度计算与动态算子融合,该方法在单个消费级设备上实现4K/8K图像的毫秒级推理,同时在多个恢复指标上超越现有最先进模型。
原文摘要 · Abstract (English)
Considering efficiency, ultra-high-definition (UHD) low-light image restoration is extremely challenging. Existing methods based on Transformer architectures or high-dimensional complex convolutional neural networks often suffer from the "memory wall" bottleneck, failing to achieve millisecond-level inference on edge devices. To address this issue, we propose a novel real-time UHD low-light enhancement network based on geometric feature fusion using Clifford algebra in 2D Euclidean space. First, we construct a four-layer feature pyramid with gradually increasing resolution, which decomposes input images into low-frequency and high-frequency structural components via a Gaussian blur kernel, and adopts a lightweight U-Net based on depthwise separable convolution for dual-branch feature extraction. Second, to resolve structural information loss and artifacts from traditional high-low frequency feature fusion, we introduce spatially aware Clifford algebra, which maps feature tensors to a multivector space (scalars, vectors, bivectors) and uses Clifford similarity to aggregate features while suppressing noise and preserving textures. In the reconstruction stage, the network outputs adaptive Gamma and Gain maps, which perform physically constrained non-linear brightness adjustment via Retinex theory. Integrated with FP16 mixed-precision computation and dynamic operator fusion, our method achieves millisecond-level inference for 4K/8K images on a single consumer-grade device, while outperforming state-of-the-art (SOTA) models on several restoration metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。