arXiv:2602.08699cs.CV2026-02

提出视图感知的低光视频增强框架,提升动态场景还原效果。

Low-Light Video Enhancement with An Effective Spatial-Temporal Decomposition Paradigm

  • 分离视图无关与视图相关成分,结合跨帧对应关系建模
  • 引入双结构网络实现帧间特征一致,参数开销小
  • 支持双向学习,适合高动态或真实复杂场景

低光视频增强(LLVE)旨在恢复因严重光照不足和噪声导致不可见的动态或静态画面。本文提出一种创新的视频分解策略,将图像分解为视图无关和视图相关成分,构建视图感知低光视频增强(VLLVE)框架。通过利用跨帧动态对应关系建模视图无关项(主要捕捉固有外观),并施加场景级连续性约束于视图相关项(主要描述光照条件),实现一致且高质量的分解。为确保一致性,设计双结构增强网络,引入跨帧交互机制,通过同时监督多帧,促使各帧具备匹配的分解特征,可无缝集成至编码器-解码器单帧网络,附加参数极少。在此基础上,进一步提出更全面的分解策略VLLVE++,引入可加残差项以模拟难以用分解建模的场景自适应退化,显著提升整体内容表征能力。此外,VLLVE++支持端到端的双向学习,同时优化增强结果与退化感知的对应关系,有效提升可靠对应,过滤错误匹配。大量实验在主流LLVE基准上验证了其优越性,尤其在真实场景及高动态视频中表现突出。

原文摘要 · Abstract (English)

Low-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition strategy that incorporates view-independent and view-dependent components to enhance the performance of LLVE. The framework is called View-aware Low-light Video Enhancement (VLLVE). We leverage dynamic cross-frame correspondences for the view-independent term (which primarily captures intrinsic appearance) and impose a scene-level continuity constraint on the view-dependent term (which mainly describes the shading condition) to achieve consistent and satisfactory decomposition results. To further ensure consistent decomposition, we introduce a dual-structure enhancement network featuring a cross-frame interaction mechanism. By supervising different frames simultaneously, this network encourages them to exhibit matching decomposition features. This mechanism can seamlessly integrate with encoder-decoder single-frame networks, incurring minimal additional parameter costs. Building upon VLLVE, we propose a more comprehensive decomposition strategy by introducing an additive residual term, resulting in VLLVE++. This residual term can simulate scene-adaptive degradations, which are difficult to model using a decomposition formulation for common scenes, thereby further enhancing the ability to capture the overall content of videos. In addition, VLLVE++ enables bidirectional learning for both enhancement and degradation-aware correspondence refinement (end-to-end manner), effectively increasing reliable correspondences while filtering out incorrect ones. Notably, VLLVE++ demonstrates strong capability in handling challenging cases, such as real-world scenes and videos with high dynamics. Extensive experiments are conducted on widely recognized LLVE benchmarks.

低光增强视频处理分解建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。