arXiv:2507.06744cs.CVcs.LG2025-07被引 6

解决文本到人像匹配中一对多身份关联难题

Dual-Granularity Cross-Modal Identity Association for Weakly-Supervised Text-to-Person Image Matching

  • 提出局部-全局双粒度身份关联机制
  • 在多个数据集上显著提升匹配准确率
  • 适合弱监督图像检索与身份识别场景

弱监督文本到人像图像匹配作为降低模型对大规模人工标注样本依赖的关键方法,具有重要研究价值。然而,现有方法难以预测复杂的多对一身份关系,严重制约性能提升。为此,本文提出一种局部-全局双粒度身份关联机制:在局部层面,显式建立批次内跨模态身份关系,强化不同模态间身份约束,更好捕捉细微差异与相关性;在全局层面,以视觉模态为锚点构建动态跨模态身份关联网络,并引入基于置信度的动态调整机制,有效提升模型对弱关联样本的识别能力与整体敏感性。此外,提出信息不对称样本对构造方法结合一致性学习,解决困难样本挖掘问题,增强模型鲁棒性。实验结果表明,所提方法显著提升跨模态匹配准确率,为文本到人像图像匹配提供高效实用解决方案。

原文摘要 · Abstract (English)

Weakly supervised text-to-person image matching, as a crucial approach to reducing models' reliance on large-scale manually labeled samples, holds significant research value. However, existing methods struggle to predict complex one-to-many identity relationships, severely limiting performance improvements. To address this challenge, we propose a local-and-global dual-granularity identity association mechanism. Specifically, at the local level, we explicitly establish cross-modal identity relationships within a batch, reinforcing identity constraints across different modalities and enabling the model to better capture subtle differences and correlations. At the global level, we construct a dynamic cross-modal identity association network with the visual modality as the anchor and introduce a confidence-based dynamic adjustment mechanism, effectively enhancing the model's ability to identify weakly associated samples while improving overall sensitivity. Additionally, we propose an information-asymmetric sample pair construction method combined with consistency learning to tackle hard sample mining and enhance model robustness. Experimental results demonstrate that the proposed method substantially boosts cross-modal matching accuracy, providing an efficient and practical solution for text-to-person image matching.

跨模态匹配弱监督学习身份识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。