arXiv:2503.10324cs.CVcs.MM2025-03CVPR被引 37

用倒置文本增强多模态特征,提升物体重识别精度

IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification

  • 通过倒置文本引导融合视觉与语义信息
  • 在三个新基准上实现更优的重识别性能
  • 适合需要跨模态检索的智能监控场景

多模态物体重识别旨在利用不同模态的互补信息检索特定物体。现有方法主要聚焦于异构视觉特征的融合,忽略了文本语义信息的潜力。为此,我们构建了三个文本增强的多模态物体重识别基准,提出基于多模态大模型(MLLMs)的标准化多模态标题生成流程,实现结构化、简洁的文本标注。此外,当前方法常直接聚合多模态信息,未选择代表性局部特征,导致冗余和高复杂度。为此,我们提出IDEA框架,包含倒置多模态特征提取器(IMFE)与协同可变形聚合(CDA)。IMFE利用模态前缀与Inversenet,结合倒置文本的语义指导实现多模态信息融合;CDA自适应生成采样位置,使模型关注全局特征与判别性局部特征的交互。在三个多模态物体重识别基准上的大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Multi-modal object Re-IDentification (ReID) aims to retrieve specific objects by utilizing complementary information from various modalities. However, existing methods focus on fusing heterogeneous visual features, neglecting the potential benefits of text-based semantic information. To address this issue, we first construct three text-enhanced multi-modal object ReID benchmarks. To be specific, we propose a standardized multi-modal caption generation pipeline for structured and concise text annotations with Multi-modal Large Language Models (MLLMs). Besides, current methods often directly aggregate multi-modal information without selecting representative local features, leading to redundancy and high complexity. To address the above issues, we introduce IDEA, a novel feature learning framework comprising the Inverted Multi-modal Feature Extractor (IMFE) and Cooperative Deformable Aggregation (CDA). The IMFE utilizes Modal Prefixes and an InverseNet to integrate multi-modal information with semantic guidance from inverted text. The CDA adaptively generates sampling positions, enabling the model to focus on the interplay between global features and discriminative local features. With the constructed benchmarks and the proposed modules, our framework can generate more robust multi-modal features under complex scenarios. Extensive experiments on three multi-modal object ReID benchmarks demonstrate the effectiveness of our proposed method.

多模态重识别文本增强特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。