通过分层匹配与位置渐进修正,提升视觉定位精度。
Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction for Visual Grounding
- 将句子分块后分层匹配图文特征
- 逐级优化目标框位置,提升定位准确率
- 适合需要精细定位的视觉理解任务
视觉定位因在多种视觉语言任务中的广泛应用而受到广泛关注。尽管该领域已取得显著进展,现有方法忽视了文本与图像特征在不同层次上的关联对跨模态匹配的促进作用。本文提出一种短语解耦跨模态分层匹配与渐进位置修正的视觉定位方法。首先通过解耦的句子短语生成掩码,构建文本与图像的分层匹配机制,凸显多层次关联在跨模态匹配中的作用。此外,基于该分层匹配机制设计目标物体位置的渐进修正策略,随着文本描述确定性提升,持续优化目标框位置。该设计探索了不同层次特征间的关联,突出了与目标物体及其位置相关的特征在定位中的关键作用。通过在多个数据集上的实验验证,该方法在性能上优于当前最先进方法。
原文摘要 · Abstract (English)
Visual grounding has attracted wide attention thanks to its broad application in various visual language tasks. Although visual grounding has made significant research progress, existing methods ignore the promotion effect of the association between text and image features at different hierarchies on cross-modal matching. This paper proposes a Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction Visual Grounding method. It first generates a mask through decoupled sentence phrases, and a text and image hierarchical matching mechanism is constructed, highlighting the role of association between different hierarchies in cross-modal matching. In addition, a corresponding target object position progressive correction strategy is defined based on the hierarchical matching mechanism to achieve accurate positioning for the target object described in the text. This method can continuously optimize and adjust the bounding box position of the target object as the certainty of the text description of the target object improves. This design explores the association between features at different hierarchies and highlights the role of features related to the target object and its position in target positioning. The proposed method is validated on different datasets through experiments, and its superiority is verified by the performance comparison with the state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。