用期望引导对齐,提升开放词汇定位的精度与效率。
ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding
- 基于软MIL池化实现隐式标记与实例选择,无需额外标注。
- 在LVIS数据集上达到36.2 AP_r,长尾类别表现显著提升。
- 轻量高效,适合零样本实例分割与检测场景。
开放词汇定位需在弱监督下实现精准视觉-语言对齐,现有方法或依赖缺乏细粒度表达能力的全局句向量,或引入需要显式标注或复杂交叉注意力的设计。本文提出ExpAlign,一种基于理论支撑的多实例学习框架。其引入期望对齐头,通过基于注意力的软MIL池化对词-区域相似性进行聚合,实现隐式词与实例选择,无需额外标注。为进一步稳定对齐学习,设计了基于能量的多尺度一致性正则化方案,包含Top-K多正例对比目标与由拉格朗日约束自由能最小化推导出的几何感知一致性目标。大量实验表明,ExpAlign在开放词汇检测与零样本实例分割任务中持续提升性能,尤其在长尾类别上表现突出。最显著的是,在LVIS minival划分上达到36.2 AP_r,优于其他同类先进方法,且模型轻量、推理高效。
原文摘要 · Abstract (English)
Open-vocabulary grounding requires accurate vision-language alignment under weak supervision, yet existing methods either rely on global sentence embeddings that lack fine-grained expressiveness or introduce token-level alignment with explicit supervision or heavy cross-attention designs. We propose ExpAlign, a theoretically grounded vision-language alignment framework built on a principled multiple instance learning formulation. ExpAlign introduces an Expectation Alignment Head that performs attention-based soft MIL pooling over token-region similarities, enabling implicit token and instance selection without additional annotations. To further stabilize alignment learning, we develop an energy-based multi-scale consistency regularization scheme, including a Top-K multi-positive contrastive objective and a Geometry-Aware Consistency Objective derived from a Lagrangian-constrained free-energy minimization. Extensive experiments show that ExpAlign consistently improves open-vocabulary detection and zero-shot instance segmentation, particularly on long-tail categories. Most notably, it achieves 36.2 AP$_r$ on the LVIS minival split, outperforming other state-of-the-art methods at comparable model scale, while remaining lightweight and inference-efficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。