用小波域掩码建模提升HDR视频的色彩一致性和时间连贯性
Wavelet-Domain Masked Image Modeling for Color-Consistent HDR Video Reconstruction
- 在小波域通过掩码建模预训练,增强色彩恢复能力
- 引入时序专家混合与动态记忆模块,显著减少闪烁并保持细节
- 新构建的场景级数据集助力性能评估,适合视频重建研究者
高动态范围(HDR)视频重建旨在从低动态范围(LDR)视频中恢复精细亮度、色彩和细节。现有方法常面临色彩失真和时序不一致问题。为此,本文提出WMNet,一种基于小波域掩码图像建模(W-MIM)的新颖HDR视频重建网络。WMNet采用两阶段训练策略:第一阶段,通过在小波域有选择地遮蔽颜色和细节信息进行自重建预训练,使网络具备鲁棒的色彩修复能力;课程学习进一步优化重建过程。第二阶段,使用预训练权重微调模型以提升最终重建质量。为改善时序一致性,引入时序专家混合(T-MoE)模块和动态记忆模块(DMM)。T-MoE 自适应融合相邻帧以减少闪烁伪影,而DMM捕捉长程依赖,确保运动平滑并保留细微细节。此外,由于现有HDR视频数据集缺乏场景级分割,本文将HDRTV4K重构为HDRTV4K-Scene,建立新的基准。大量实验表明,WMNet在多个评估指标上达到最先进水平,显著提升色彩保真度、时序连贯性和感知质量。代码已公开于:https://github.com/eezkni/WMNet
原文摘要 · Abstract (English)
High Dynamic Range (HDR) video reconstruction aims to recover fine brightness, color, and details from Low Dynamic Range (LDR) videos. However, existing methods often suffer from color inaccuracies and temporal inconsistencies. To address these challenges, we propose WMNet, a novel HDR video reconstruction network that leverages Wavelet domain Masked Image Modeling (W-MIM). WMNet adopts a two-phase training strategy: In Phase I, W-MIM performs self-reconstruction pre-training by selectively masking color and detail information in the wavelet domain, enabling the network to develop robust color restoration capabilities. A curriculum learning scheme further refines the reconstruction process. Phase II fine-tunes the model using the pre-trained weights to improve the final reconstruction quality. To improve temporal consistency, we introduce the Temporal Mixture of Experts (T-MoE) module and the Dynamic Memory Module (DMM). T-MoE adaptively fuses adjacent frames to reduce flickering artifacts, while DMM captures long-range dependencies, ensuring smooth motion and preservation of fine details. Additionally, since existing HDR video datasets lack scene-based segmentation, we reorganize HDRTV4K into HDRTV4K-Scene, establishing a new benchmark for HDR video reconstruction. Extensive experiments demonstrate that WMNet achieves state-of-the-art performance across multiple evaluation metrics, significantly improving color fidelity, temporal coherence, and perceptual quality. The code is available at: https://github.com/eezkni/WMNet
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。