针对伪装场景的图文检索,提出专家协作网络提升跨模态对齐能力。
Camouflage-aware Image-Text Retrieval via Expert Collaboration
- 双分支视觉编码器:一个捕捉整体图像特征,另一个专注伪装物体表征。
- 在10.5K样本数据集上,相比主流模型提升近29%检索准确率。
- 适合研究伪装目标识别、跨模态理解与图文匹配的学者使用。
由于其广泛的实际应用价值,伪装场景理解(CSU)受到广泛关注。然而,在该领域中,鲁棒的图像-文本跨模态对齐仍缺乏深入探索,阻碍了对伪装场景及其相关应用的深入理解。为此,我们聚焦典型的图文检索任务,提出一项新任务——伪装感知图文检索(CA-ITR)。我们构建了一个专用的伪装图像-文本检索数据集(CamoIT),包含约10.5K个样本,配有细粒度文本标注。在该数据集上的基准测试表明,现有先进检索技术在处理CA-ITR任务时面临显著挑战,主要源于物体的伪装特性及复杂图像内容。为此,我们提出伪装专家协作网络(CECNet),其采用双分支视觉编码器:一枝捕获整体图像表示,另一枝引入专门模型注入伪装物体表征。通过新颖的置信度条件图注意力机制(C²GA),有效利用两分支间的互补性。对比实验显示,CECNet在整体检索准确率上提升约29%,超越七种代表性检索模型。数据集与代码将公开于https://github.com/jiangyao-scu/CA-ITR。
原文摘要 · Abstract (English)
Camouflaged scene understanding (CSU) has attracted significant attention due to its broad practical implications. However, in this field, robust image-text cross-modal alignment remains under-explored, hindering deeper understanding of camouflaged scenarios and their related applications. To this end, we focus on the typical image-text retrieval task, and formulate a new task dubbed ``camouflage-aware image-text retrieval'' (CA-ITR). We first construct a dedicated camouflage image-text retrieval dataset (CamoIT), comprising $\sim$10.5K samples with multi-granularity textual annotations. Benchmark results conducted on CamoIT reveal the underlying challenges of CA-ITR for existing cutting-edge retrieval techniques, which are mainly caused by objects' camouflage properties as well as those complex image contents. As a solution, we propose a camouflage-expert collaborative network (CECNet), which features a dual-branch visual encoder: one branch captures holistic image representations, while the other incorporates a dedicated model to inject representations of camouflaged objects. A novel confidence-conditioned graph attention (C\textsuperscript{2}GA) mechanism is incorporated to exploit the complementarity across branches. Comparative experiments show that CECNet achieves $\sim$29% overall CA-ITR accuracy boost, surpassing seven representative retrieval models. The dataset and code will be available at https://github.com/jiangyao-scu/CA-ITR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。