提出首个超高清多焦点图像融合数据集与轻量级查询表框架,实现4K实时融合。
UHD-MFF: Shattering Barriers in Multi-Focus Ultra-High-Definition Image Fusion via Learnable Lookup Tables

- 设计分层查询表结构,低分辨率处理区域决策,高分辨率捕捉边缘细节。
- 在4K分辨率下实现实时融合,计算开销极低,支持手机端部署。
- 首次构建超高清多焦点融合数据集,解决数据稀缺难题,推动应用落地。
随着成像技术进步,超高清图像在现代视觉应用中愈发重要。然而,现有多焦点图像融合方法仍局限于低分辨率图像,在超高清场景下面临数据可用性、模型适应性和部署可行性三大障碍,严重制约实际应用。为此,本文首先构建了首个大规模超高清多焦点融合数据集UHD-MFF;其次提出专为超高清图像设计的可学习查找表框架UMF-LUT,包含粗粒度区域查找表(C-LUT)和细节边缘查找表(D-LUT)。C-LUT在低分辨率尺度联合查询梯度与语义线索,实现区域级决策;D-LUT在高分辨率尺度利用高效拉普拉斯线索提供边缘级补充信息。该设计使模型特别适用于超高清多焦点图像融合。最后,该框架具备强部署性,计算开销极小,支持4K实时融合,展现出在智能手机上的应用潜力。大量实验表明,其在视觉保真度和定量指标上均优于当前最优方法,有效推动多焦点图像融合向超高清场景发展。代码已开源:https://github.com/zyb5/UHD-MFF。
原文摘要 · Abstract (English)
With the advancement of imaging technology, ultra-high-definition images have become increasingly essential in modern visual applications. However, existing multi-focus image fusion remains largely confined to low-resolution images and faces three major barriers in UHD scenarios, namely data availability, model adaptability, and deployment feasibility, which severely hinder its practical application. To shatter these barriers, first, we propose the UHD-MFF dataset, the first large-scale ultra-high-resolution multi-focus fusion dataset. Second, we propose a scale-specialized lookup-table framework tailored for ultra-high-resolution images, termed as UMF-LUT. It consists of Coarse-Region Lookup Table (C-LUT) and Detail-Edge Lookup Table (D-LUT). Specifically, C-LUT performs joint queries of multiple gradient cues and semantic cues at low-resolution scales to enable region-level decision-making. Also, D-LUT operates at high-resolution scales, leveraging efficient Laplacian cues to provide complementary edge-level decision information. Such a design makes the model particularly well-suited for ultra-high-resolution multi-focus image fusion. Finally, it offers strong deployability with minimal computational overhead, enabling real-time 4K multi-focus fusion and showing promising potential for smartphone. Extensive experiments demonstrate that it outperforms SOTA methods in both visual fidelity and quantitative metrics. It effectively advances the development of multi-focus image fusion toward ultra-high-resolution imaging scenarios. The code is available at https://github.com/zyb5/UHD-MFF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。