提升暗光视频实例分割效果,让模型在低光照下更准更稳。
ELVIS: Enhance Low-Light for Video Instance Segmentation in the Dark
- 构建无监督合成暗光视频数据流,模拟空间与时间退化。
- 提升合成数据上性能达+3.7AP,真实暗光视频超基线+2.8AP。
- 无需校准,自动估计退化特征,适合夜间监控等场景应用。
低光照视频实例分割(VIS)对人类和机器都极具挑战,受噪声、模糊等退化影响严重。现有大规模标注数据集匮乏,且合成管道难以建模时序退化,制约了进展。同时,现有VIS方法对低光照退化不鲁棒,微调后仍表现不佳。本文提出ELVIS(Enhance Low-Light for Video Instance Segmentation),实现先进VIS模型在暗光场景的域自适应。ELVIS包含无监督合成暗光视频流水线(建模空间与时间退化)、免校准退化特征估计网络(VDP-Net)及解耦退化的增强解码头。在合成暗光YouTube-VIS 2019数据集上性能提升最高达+3.7AP,真实暗光视频上优于两阶段基线至少+2.8AP。代码与数据集见:https://joannelin168.github.io/research/ELVIS
原文摘要 · Abstract (English)
Video instance segmentation (VIS) for low-light content remains highly challenging for both humans and machines alike, due to noise, blur and other adverse conditions. The lack of large-scale annotated datasets and the limitations of current synthetic pipelines, particularly in modeling temporal degradations, further hinder progress. Moreover, existing VIS methods are not robust to the degradations found in low-light videos and, consequently, perform poorly even after finetuning. In this paper, we introduce \textbf{ELVIS} (\textbf{E}nhance \textbf{L}ow-Light for \textbf{V}ideo \textbf{I}nstance \textbf{S}egmentation), a framework that enables domain adaptation of state-of-the-art VIS models to low-light scenarios. ELVIS is comprised of an unsupervised synthetic low-light video pipeline that models both spatial and temporal degradations, a calibration-free degradation profile estimation network (VDP-Net) and an enhancement decoder head that disentangles degradations from content features. ELVIS improves performances by up to \textbf{+3.7AP} on the synthetic low-light YouTube-VIS 2019 dataset and beats two-stage baselines by at least \textbf{+2.8AP} on real low-light videos. Code and dataset available at: \href{https://joannelin168.github.io/research/ELVIS}{https://joannelin168.github.io/research/ELVIS}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。