解决视觉定位中的模态偏差与语义理解不足问题
BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual Grounding
- 分离模态特征,通过三模块增强指代表达理解
- 在五个数据集上达到最优性能,且计算更高效
- 适合关注多模态推理与公平性的研究者
视觉定位(VG)旨在定位由表达式所指的特定区域,是多模态理解中的基础但具挑战性任务。尽管近期的一塔架构迁移方法推动了该领域进展,但仍面临两大局限:(1) 多模态表征过度纠缠,加剧了误导性模态偏差;(2) 语义推理不足,影响对指代线索的理解。本文提出BARE框架,一种面向偏差感知与推理增强的一塔视觉定位方法。BARE引入新机制以保留模态特异性特征,并通过三个创新模块构建指代表意:(i) 语言显著性调制器,(ii) 视觉偏差校正,(iii) 指代关系增强,共同缓解多模态干扰并提升指代理解。在五个基准上的大量实验表明,BARE不仅达到当前最佳性能,还相较现有方法具有更优计算效率。代码已公开于https://github.com/Marloweeee/BARE。
原文摘要 · Abstract (English)
Visual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent grounding transfer works have advanced the field through one-tower architectures, they still suffer from two primary limitations: (1) over-entangled multimodal representations that exacerbate deceptive modality biases, and (2) insufficient semantic reasoning that hinders the comprehension of referential cues. In this paper, we propose BARE, a bias-aware and reasoning-enhanced framework for one-tower visual grounding. BARE introduces a mechanism that preserves modality-specific features and constructs referential semantics through three novel modules: (i) language salience modulator, (ii) visual bias correction and (iii) referential relationship enhancement, which jointly mitigate multimodal distractions and enhance referential comprehension. Extensive experimental results on five benchmarks demonstrate that BARE not only achieves state-of-the-art performance but also delivers superior computational efficiency compared to existing approaches. The code is publicly accessible at https://github.com/Marloweeee/BARE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。