解决视频去摩尔纹中空间变化、大尺度结构与通道依赖问题。
MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing
- 分步解耦摩尔纹与内容,动态生成适应性修复特征。
- 在RAW和sRGB数据集上,PSNR、SSIM、LPIPS全面领先。
- 无需显式对齐,隐式保持帧间时序一致性,适合移动拍摄场景。
便携成像技术的普及使基于相机的屏幕捕获成为常态,但摄像头色彩滤波阵列(CFA)与显示屏子像素间的频率混叠会导致严重摩尔纹,损害图像质量。现有去摩尔纹方法仍面临四大挑战:(i) 帧内摩尔纹强度空间变化,(ii) 大范围全局结构,(iii) 通道相关性统计,(iv) 帧间快速时序波动。为此,本文提出摩尔纹条件混合自适应变换器(MoCHA-former),包含两个核心组件:解耦摩尔纹自适应去摩尔纹(DMAD)与时空自适应去摩尔纹(STAD)。DMAD通过摩尔纹解耦块(MDB)与细节解耦块(DDB)分离摩尔纹与内容,并利用摩尔纹条件块(MCB)生成自适应特征以实现精准修复;STAD引入空间融合块(SFB)结合窗口注意力捕捉大尺度结构,以及特征通道注意力(FCA)建模RAW帧中的通道依赖关系。为保证时序一致性,模型采用隐式帧对齐,无需显式对齐模块。通过定性和定量分析摩尔纹特性,在两个涵盖RAW与sRGB域的视频数据集上评估,结果表明MoCHA-former在PSNR、SSIM与LPIPS指标上持续优于已有方法。
原文摘要 · Abstract (English)
Recent advances in portable imaging have made camera-based screen capture ubiquitous. Unfortunately, frequency aliasing between the camera's color filter array (CFA) and the display's sub-pixels induces moiré patterns that severely degrade captured photos and videos. Although various demoiréing models have been proposed to remove such moiré patterns, these approaches still suffer from several limitations: (i) spatially varying artifact strength within a frame, (ii) large-scale and globally spreading structures, (iii) channel-dependent statistics and (iv) rapid temporal fluctuations across frames. We address these issues with the Moiré Conditioned Hybrid Adaptive Transformer (MoCHA-former), which comprises two key components: Decoupled Moiré Adaptive Demoiréing (DMAD) and Spatio-Temporal Adaptive Demoiréing (STAD). DMAD separates moiré and content via a Moiré Decoupling Block (MDB) and a Detail Decoupling Block (DDB), then produces moiré-adaptive features using a Moiré Conditioning Block (MCB) for targeted restoration. STAD introduces a Spatial Fusion Block (SFB) with window attention to capture large-scale structures, and a Feature Channel Attention (FCA) to model channel dependence in RAW frames. To ensure temporal consistency, MoCHA-former performs implicit frame alignment without any explicit alignment module. We analyze moiré characteristics through qualitative and quantitative studies, and evaluate on two video datasets covering RAW and sRGB domains. MoCHA-former consistently surpasses prior methods across PSNR, SSIM, and LPIPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。