提出跨模态统一网络,让不同模态数据共享通用特征,同时保留身份信息。
Mind the Gap: Learning Modality-Agnostic Representations with a Cross-Modality UNet

- 用交叉模态变换与同模态重建学习通用表示
- 在5个任务中超越现有方法,尤其在光谱与人脸匹配上表现突出
- 抗遮挡能力可反映模型是否真正解决模态差异问题
跨模态识别在科学、执法和娱乐中具有重要应用。现有方法或消除模态特异性导致判别信息丢失,或依赖模态转换但易失效。本文提出紧凑的编码器-解码器模块cmUNet,通过跨模态变换与同模态重建,在保留身份相关特征的同时学习模态无关表示,并引入对抗性/感知损失增强原始空间表示的不可区分性。针对跨模态匹配,设计MarrNet,将cmUNet与标准特征提取网络结合,输出匹配相似度。在五项挑战性任务中验证:拉曼-红外光谱匹配、跨模态行人重识别及异质人脸识别(照片-素描、可见-近红外、可见-热成像),MarrNet性能优于现有最佳方法。此外发现,部分方法因无法有效处理模态差异,会偏向提取局部甚至错误区域的判别特征,导致泛化能力差;本文表明,对遮挡的鲁棒性可作为判断方法是否真正跨越模态鸿沟的重要指标。
原文摘要 · Abstract (English)
Cross-modality recognition has many important applications in science, law enforcement and entertainment. Popular methods to bridge the modality gap include reducing the distributional differences of representations of different modalities, learning indistinguishable representations or explicit modality transfer. The first two approaches suffer from the loss of discriminant information while removing the modality-specific variations. The third one heavily relies on the successful modality transfer, could face catastrophic performance drop when explicit modality transfers are not possible or difficult. To tackle this problem, we proposed a compact encoder-decoder neural module (cmUNet) to learn modality-agnostic representations while retaining identity-related information. This is achieved through cross-modality transformation and in-modality reconstruction, enhanced by an adversarial/perceptual loss which encourages indistinguishability of representations in the original sample space. For cross-modality matching, we propose MarrNet where cmUNet is connected to a standard feature extraction network which takes as inputs the modality-agnostic representations and outputs similarity scores for matching. We validated our method on five challenging tasks, namely Raman-infrared spectrum matching, cross-modality person re-identification and heterogeneous (photo-sketch, visible-near infrared and visible-thermal) face recognition, where MarrNet showed superior performance compared to state-of-the-art methods. Furthermore, it is observed that a cross-modality matching method could be biased to extract discriminant information from partial or even wrong regions, due to incompetence of dealing with modality gaps, which subsequently leads to poor generalization. We show that robustness to occlusions can be an indicator of whether a method can well bridge the modality gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。