通过稀疏解耦与文本图像差异过滤,提升多模态物体重识别的准确率与可解释性。
MODAL: Multi-Modal Object Re-ID via Model-Driven Sparse Decoupling and Text-Image Differential Filtering

- 基于耦合稀疏编码理论,分离视觉与文本特征中的单模、双模及三模共享成分。
- 在缺失模态情况下仍保持性能,因仅激活一致共享子空间。
- 利用粗粒度文本语义抑制无关视觉响应,增强判别性特征表达。
多模态物体重识别(Re-ID)旨在通过融合视觉(如RGB、NIR、TIR)与文本模态的互补信息,在复杂环境中实现跨摄像头物体检索。然而,现有方法缺乏合理的特征解耦与一致的多模态融合,导致表示纠缠,引发跨模态冲突、判别性线索模糊,并在模态缺失时出现分布偏移。为此,本文提出MODAL框架,基于耦合稀疏编码理论与差分抑制原理。核心是多模态特征稀疏解耦模块,以模型驱动的深度展开方式构建,显式分解多模态特征为单模特定、双模与三模共享表示,实现更透明有效的特征解耦。得益于这一解耦机制,MODAL通过模态感知子空间激活,在不完整模态场景下自然缓解性能下降。此外,提出文本-图像差异过滤模块,利用粗粒度文本语义自适应抑制解耦后视觉表示中的任务无关响应,从而增强判别性信息。在四个数据集上的大量实验表明,MODAL达到当前最优性能,并具有优异的透明性。
原文摘要 · Abstract (English)
Multi-modal object re-identification (Re-ID) aims to facilitate cross-camera object retrieval in complex environments by leveraging complementary information from visual (e.g., RGB, NIR, TIR) and textual modalities. However, existing approaches often lack principled feature disentanglement and coherent multi-modal integration, leading to entangled representations that introduce cross-modal conflicts, obscure discriminative cues, and suffer distribution shift under modality-missing conditions. To tackle these challenges, we propose MODAL, a novel multi-modal object re-identification framework, grounded in coupled sparse coding theory and differential suppression principles. A core component of MODAL is a Multi-modal Feature Sparse Decoupling module, developed in a model-driven deep unrolling manner based on multi-modal coupled sparse coding. It explicitly decomposes multi-modal features into uni-modal specific, bi-modal and tri-modal shared representations, thereby achieving more transparent and effective feature disentanglement. Benefiting from the principled feature disentanglement, MODAL naturally mitigates performance degradation in incomplete-modality scenarios via a Modality-Aware Subspace Activation that selectively activates only the consistently shared subspaces. Moreover, we propose a Text-Image Differential Filtering module that leverages coarse-grained textual semantics to adaptively suppress task-irrelevant responses in the decoupled visual representations, thereby enhancing discriminative information. Extensive experiments on four datasets demonstrate that MODAL achieves state-of-the-art performance with superior transparency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。