arXiv:2504.16145cs.CVcs.AI2025-04被引 4

通过渐进式语言引导提升多任务视觉定位性能,无需额外融合模块。

Progressive Language-guided Visual Learning for Multi-Task Visual Grounding

  • 渐进注入语言信息到视觉主干,增强特征表达
  • 在多个基准数据集上显著超越现有方法
  • 共享任务头促进两任务协同,适合多任务视觉理解研究

多任务视觉定位(MTVG)包含指代表达理解(REC)和指代表达分割(RES)两个子任务。现有方法通常采用三阶段流程:分别提取视觉与语言模态特征、跨模态交互模块、独立预测头。尽管表现优异,但仍存在两大问题:1)语言信息未充分注入视觉主干,需额外跨模态交互模块;2)未能有效利用REC与RES之间的关系以实现协同预测。为此,本文提出渐进式语言引导视觉学习框架PLVL,不仅挖掘视觉模态自身特征表达,还逐步注入语言信息以学习语言相关视觉特征,从而无需额外跨模态融合模块即可实现全面语言引导。进一步分析发现,REC的定位中心可辅助识别RES的目标区域,据此设计多任务头实现协同预测。在多个基准数据集上的大量实验表明,PLVL在REC与RES任务中均显著优于代表性方法。

原文摘要 · Abstract (English)

Multi-task visual grounding (MTVG) includes two sub-tasks, i.e., Referring Expression Comprehension (REC) and Referring Expression Segmentation (RES). The existing representative approaches generally follow the research pipeline which mainly consists of three core procedures, including independent feature extraction for visual and linguistic modalities, respectively, cross-modal interaction module, and independent prediction heads for different sub-tasks. Albeit achieving remarkable performance, this research line has two limitations: 1) The linguistic content has not been fully injected into the entire visual backbone for boosting more effective visual feature extraction and it needs an extra cross-modal interaction module; 2) The relationship between REC and RES tasks is not effectively exploited to help the collaborative prediction for more accurate output. To deal with these problems, in this paper, we propose a Progressive Language-guided Visual Learning framework for multi-task visual grounding, called PLVL, which not only finely mine the inherent feature expression of the visual modality itself but also progressively inject the language information to help learn linguistic-related visual features. In this manner, our PLVL does not need additional cross-modal fusion module while fully introducing the language guidance. Furthermore, we analyze that the localization center for REC would help identify the to-be-segmented object region for RES to some extent. Inspired by this investigation, we design a multi-task head to accomplish collaborative predictions for these two sub-tasks. Extensive experiments conducted on several benchmark datasets comprehensively substantiate that our PLVL obviously outperforms the representative methods in both REC and RES tasks. https://github.com/jcwang0602/PLVL

视觉定位多任务学习语言引导协同预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。