用不准确的描述也能去镜面,模型自适应调整语言信息。
Adaptive Language-Aware Image Reflection Removal Network
- 自适应融合语言与视觉特征,动态过滤干扰信息
- 在语言描述错误时仍能保持良好去反射效果
- 适合处理模糊扭曲的复杂反射场景
现有图像去反射方法难以应对复杂反射。准确的语言描述有助于模型理解内容,但反射图像中的模糊和扭曲导致机器生成的语言描述常不准确,影响语言引导去反射性能。为此,我们提出自适应语言感知网络(ALANet),即使在语言输入不准确时也能有效去反射。ALANet结合滤波与优化策略:滤波策略降低语言干扰,保留其优势;优化策略增强语言与视觉特征的对齐。同时利用语言提示解耦特征图中特定层内容,提升处理复杂反射的能力。为评估模型在复杂反射及不同语言准确率下的表现,我们构建了复杂反射与语言准确率变化数据集(CRLAV)。实验表明,ALANet优于当前最优方法。代码与数据集已公开于 https://github.com/fashyon/ALANet。
原文摘要 · Abstract (English)
Existing image reflection removal methods struggle to handle complex reflections. Accurate language descriptions can help the model understand the image content to remove complex reflections. However, due to blurred and distorted interferences in reflected images, machine-generated language descriptions of the image content are often inaccurate, which harms the performance of language-guided reflection removal. To address this, we propose the Adaptive Language-Aware Network (ALANet) to remove reflections even with inaccurate language inputs. Specifically, ALANet integrates both filtering and optimization strategies. The filtering strategy reduces the negative effects of language while preserving its benefits, whereas the optimization strategy enhances the alignment between language and visual features. ALANet also utilizes language cues to decouple specific layer content from feature maps, improving its ability to handle complex reflections. To evaluate the model's performance under complex reflections and varying levels of language accuracy, we introduce the Complex Reflection and Language Accuracy Variance (CRLAV) dataset. Experimental results demonstrate that ALANet surpasses state-of-the-art methods for image reflection removal. The code and dataset are available at https://github.com/fashyon/ALANet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。