用无限像素学习技术,在普通显卡上实时融合超高清动态多曝光图像。
Ultra-High-Definition Dynamic Multi-Exposure Image Fusion via Infinite Pixel Learning
- 将图像序列切片并用注意力缓存处理,实现无限长数据流的高效建模。
- 在单张消费级显卡上实现超过40fps的实时融合,保持高质量视觉效果。
- 适用于需要低延迟高画质的移动设备或嵌入式系统场景。
随着设备成像分辨率的持续提升,超高清(UHD)图像日益普及。然而,现有动态场景多曝光图像融合方法针对低分辨率图像设计,难以在资源受限设备上高效生成高质量UHD图像。受大语言模型处理无限文本的启发,我们提出一种新型学习范式——无限像素学习(IPL),可在单张消费级GPU上实现UHD动态多曝光图像的实时融合。方法包含三个核心组件:首先将输入序列切片以缓解模型处理压力;其次引入类似KV缓存的注意力缓存机制,支持无限数据流处理;最后设计缓存压缩方法,减轻设备存储负担。此外,我们构建了一个新的UHD基准测试集评估方法有效性。大量实验表明,该方法在单张消费级GPU上实现超过40fps的实时融合,同时保持优异的视觉质量。
原文摘要 · Abstract (English)
With the continuous improvement of device imaging resolution, the popularity of Ultra-High-Definition (UHD) images is increasing. Unfortunately, existing methods for fusing multi-exposure images in dynamic scenes are designed for low-resolution images, which makes them inefficient for generating high-quality UHD images on a resource-constrained device. To alleviate the limitations of extremely long-sequence inputs, inspired by the Large Language Model (LLM) for processing infinitely long texts, we propose a novel learning paradigm to achieve UHD multi-exposure dynamic scene image fusion on a single consumer-grade GPU, named Infinite Pixel Learning (IPL). The design of our approach comes from three key components: The first step is to slice the input sequences to relieve the pressure generated by the model processing the data stream; Second, we develop an attention cache technique, which is similar to KV cache for infinite data stream processing; Finally, we design a method for attention cache compression to alleviate the storage burden of the cache on the device. In addition, we provide a new UHD benchmark to evaluate the effectiveness of our method. Extensive experimental results show that our method maintains high-quality visual performance while fusing UHD dynamic multi-exposure images in real-time (>40fps) on a single consumer-grade GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。