arXiv:2505.20001cs.CV2025-05被引 6

用文本调控专家网络,提升跨模态物体重识别精度

NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-Identification

  • 分语义与结构两支路,分别捕捉细粒度外观与整体结构特征
  • 通过高质量文本调制专家,挖掘跨模态互补信息,显著降低识别错误率
  • 适合需要高精度跨模态识别的场景,如智能监控与自动驾驶

多模态物体重识别旨在跨异构模态获取完整身份特征。现有方法多依赖隐式特征融合模块,难以建模真实场景下的细粒度识别模式。得益于多模态大语言模型(MLLMs),物体外观可有效转化为描述性标题。本文提出基于属性置信度的可靠标题生成流程,显著降低MLLMs的未知识别率,提升生成文本质量。进一步提出新型重识别框架NEXT——基于文本调控的多粒度专家混合模型。具体而言,将识别任务解耦为语义与结构分支:语义分支引入文本调制语义专家(TMSE),随机采样高质量标题以调制专家,捕获语义特征并挖掘跨模态互补线索;结构分支设计上下文共享结构专家(CSSE),聚焦整体物体结构,通过软路由机制保持身份结构一致性。最后提出多粒度特征聚合(MGFA),采用统一融合策略整合多粒度专家特征,生成最终身份表示。在两个公开行人数据集和三个车辆数据集上的大量实验表明,该方法显著优于现有最先进方法。

原文摘要 · Abstract (English)

Multi-modal object Re-IDentification (ReID) aims to obtain complete identity features across heterogeneous modalities. However, most existing methods rely on implicit feature fusion modules, making it difficult to model fine-grained recognition patterns under various challenges in real world. Benefiting from the powerful Multi-modal Large Language Models (MLLMs), the object appearances are effectively translated into descriptive captions. In this paper, we propose a reliable caption generation pipeline based on attribute confidence, which significantly reduces the unknown recognition rate of MLLMs and improves the quality of generated text. Additionally, to model diverse identity patterns, we propose a novel ReID framework, named NEXT, the Multi-grained Mixture of Experts via Text-Modulation for Multi-modal Object Re-Identification. Specifically, we decouple the recognition problem into semantic and structural branches to separately capture fine-grained appearance features and coarsegrained structure features. For semantic recognition, we first propose a Text-Modulated Semantic Experts (TMSE), which randomly samples high-quality captions to modulate experts capturing semantic features and mining inter-modality complementary cues. Second, to recognize structure features, we propose a Context-Shared Structure Experts (CSSE), which focuses on the holistic object structure and maintains identity structural consistency via a soft routing mechanism. Finally, we propose a Multi-Grained Features Aggregation (MGFA), which adopts a unified fusion strategy to effectively integrate multi-grained expert features into the final identity representations. Extensive experiments on two public person datasets and three vehicle datasets demonstrate the effectiveness of our method, showing that it significantly outperforms existing state-of-the-art methods.

多模态重识别专家网络文本调制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。