arXiv:2606.11186cs.CV2026-06中稿 · ICML

提出可任意组合模态的低光视频增强方法,实测在无辅助模态时仍表现优异。

AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference

论文配图:AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference
图 1 · 摘自论文原文
  • 设计跨模态门控翻译器,生成隐式辅助特征以应对模态缺失
  • 在合成数据上预训练,提升多模态对齐能力,实现零假设推理
  • 支持任意模态组合,适合真实场景中模态不全的应用

低光视频增强(LLVE)因光照不足导致信息严重退化而极具挑战。现有融合事件流、红外图像等辅助模态的方法虽显著提升性能,但通常假设推理时模态始终可用,这在真实场景中难以实现。为此,本文提出AMNet,一种统一的多模态框架,支持灵活的模态无关推理,即辅助模态可能缺失。为应对模态缺失问题,提出空间-谱双门控翻译器,学习辅助模态与RGB输入间的对应关系,生成隐式辅助表示以支撑鲁棒增强。此外,基于仅含RGB的大型数据集,通过合成辅助模态进行大规模多模态预训练,充分挖掘跨模态关联。大量实验表明,AMNet能处理任意推理时的模态组合,在模态缺失条件下仍保持优越性能。代码与模型已开源。

原文摘要 · Abstract (English)

Low-light video enhancement (LLVE) remains a challenging task due to severe information degradation under low-illumination conditions. Recent multimodal approaches have significantly improved enhancement performance by incorporating auxiliary modalities, such as event streams and infrared images. However, these methods typically assume the availability of these modalities at inference, which is often not feasible in real-world scenarios. To solve this problem, in this work, we propose AMNet, a unified multimodal framework for LLVE, to support flexible modality-agnostic inference, where auxiliary modalities may be unavailable. To address the issue of modality absence, we introduce a Spatial-Spectral Dual-Gated Translator that learns the correspondence between auxiliary modalities and RGB inputs, producing implicit auxiliary representations to support the robust enhancement. Additionally, to fully facilitate the learning of cross-modal correspondence, we conduct large-scale multimodal pretraining based on the RGB-only dataset with synthetic auxiliary modalities. Extensive experiments demonstrate that AMNet could handle arbitrary inference-time modality combinations and exhibits superior performance for LLVE under modality absence conditions. Code and models are available on the project page.

低光增强多模态视频处理门控机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。