用九种视觉模态提升暗光图像质量,效果领先。
ModalFormer: Multimodal Transformer for Low-Light Image Enhancement
- 融合九种辅助模态信息,通过跨模态注意力机制增强图像
- 在多个基准数据集上达到当前最优性能
- 适合需要高精度图像增强的研究与应用
暗光图像增强(LLIE)因噪声、细节丢失和对比度差而具有挑战性。现有方法多依赖于RGB图像的像素级变换,忽视了多视觉模态提供的丰富上下文信息。本文提出首个大规模多模态框架ModalFormer,全面利用九种辅助模态,在多个基准数据集上实现顶尖性能。模型包含两个核心组件:用于融合多模态信息的交叉模态变压器(CM-T),以及多个专用于多模态特征重建的子网络。其中,创新的交叉模态多头自注意力机制(CM-MSA)有效整合了深度特征嵌入、分割信息、几何线索和颜色信息等模态特征,生成信息丰富的混合注意力图。实验验证了该方法在多个数据集上的优越性。预训练模型与结果已开源:https://github.com/albrateanu/ModalFormer。
原文摘要 · Abstract (English)
Low-light image enhancement (LLIE) is a fundamental yet challenging task due to the presence of noise, loss of detail, and poor contrast in images captured under insufficient lighting conditions. Recent methods often rely solely on pixel-level transformations of RGB images, neglecting the rich contextual information available from multiple visual modalities. In this paper, we present ModalFormer, the first large-scale multimodal framework for LLIE that fully exploits nine auxiliary modalities to achieve state-of-the-art performance. Our model comprises two main components: a Cross-modal Transformer (CM-T) designed to restore corrupted images while seamlessly integrating multimodal information, and multiple auxiliary subnetworks dedicated to multimodal feature reconstruction. Central to the CM-T is our novel Cross-modal Multi-headed Self-Attention mechanism (CM-MSA), which effectively fuses RGB data with modality-specific features--including deep feature embeddings, segmentation information, geometric cues, and color information--to generate information-rich hybrid attention maps. Extensive experiments on multiple benchmark datasets demonstrate ModalFormer's state-of-the-art performance in LLIE. Pre-trained models and results are made available at https://github.com/albrateanu/ModalFormer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。