arXiv:2608.01829cs.CV2026-08

一套模型同时修复4K视频的雾霾、雨天、低光和噪点,不依赖标签且速度快。

MoCRA: Mixture of Compositional Rank-1 Atoms for 4K All-in-One Video Restoration

论文配图:MoCRA: Mixture of Compositional Rank-1 Atoms for 4K All-in-One Video Restoration
图 1 · 摘自论文原文
  • 用分频带条件控制不同退化类型,让模型精准适配各自特征。
  • 在100个4K视频上达到11个基线中的最佳平均清晰度,修复速度<0.5秒。
  • 无需运动估计或成对数据,适合部署于真实场景的视频恢复任务。

真实视频常含雾、雨、暗、噪等问题,可部署的修复器需同时满足无退化标签、原生4K输出和播放稳定。现有方法分别应对,但在联合问题上失效:帧间退化判断波动、下采样代理丢失雨/噪信息、密集时序对齐不适应4K内存。为此我们构建了首个4K全任务基准UHV-4K-AIO,通过物理建模在100段4K干净视频上叠加共享深度与运动的雾、雨、传感器噪声和低光。其设计揭示了模型核心:雾与低光在降采样后仍存,而雨与噪仅存在于原生尺度。采用带匹配的组合条件机制,将条件容量、计算和监督集中在各退化对应频带。单一秩-1原子字典,每帧稀疏重构,同时驱动每片一次的粗略分支与浅层原分辨率精修分支,仅360万参数,无需光流。训练一次即可覆盖四项任务,达到十一个重训图像与视频基线的最佳平均PSNR,保持与基于光流视频模型相当的形变误差,且从不估计运动,原生4K修复耗时不足0.5秒,快于最快基线的1.7秒。

原文摘要 · Abstract (English)

Real-world video arrives hazy, rainy, dark, or noisy, and a deployable restorer faces three demands at once: no degradation label, native 4K output, and stability in playback. Existing methods answer them separately and break on the joint problem, because per-frame degradation readings flip between frames, downsampled proxies erase the rain and noise they are meant to remove, and dense temporal alignment does not fit 4K memory. No paired benchmark even poses that problem, so we build one. UHV-4K-AIO renders physically modeled haze, rain, sensor noise, and low light over the same 100 clean 4K clips with shared depth and motion, and its construction exposes the split MoCRA is built on: haze and low light survive aggressive downsampling, while rain and noise exist only at native scale. Band-matched compositional conditioning follows, spending conditioning capacity, computation, and supervision in the band where each degradation lives. One dictionary of rank-1 atoms, recomposed sparsely per frame, conditions both a once-per-clip coarse branch and a shallow native-resolution refiner, in 3.6M parameters and with no optical flow. Trained once for all four tasks, MoCRA takes the best task-mean PSNR of eleven retrained image and video baselines, holds warping error at the level of the flow-based video models while never estimating motion, and restores native 4K in under half a second, against 1.7 seconds for the fastest baseline.

视频修复4K处理多退化轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。