arXiv:2505.19242cs.CV2025-05被引 1

用视觉语言模型提升指代分割的精准度和跨模态对齐

Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model

  • 引入可变形卷积与SE块增强特征适应性
  • 提出新损失函数,提升边界精度与类别平衡
  • 适合需要精准对象定位的多模态应用

图像分割是计算机视觉的基础任务,旨在将图像划分为语义有意义的区域。指代图像分割通过自然语言表达定位特定物体,需有效融合视觉与语言信息。本文提出SegVLM,一种结合架构改进的视觉语言模型,以提升分割精度与跨模态对齐能力。模型引入挤压-激励(SE)块进行动态特征重校准,采用可变形卷积实现几何适应性,结合残差连接支持深层特征学习。此外,提出新型指代感知融合(RAF)损失,平衡区域对齐、边界精度与类别不平衡问题。大量实验与消融研究显示,各组件均带来稳定性能提升。SegVLM在多个数据集及指代表达场景下表现出强泛化能力。

原文摘要 · Abstract (English)

Image segmentation is a fundamental task in computer vision, aimed at partitioning an image into semantically meaningful regions. Referring image segmentation extends this task by using natural language expressions to localize specific objects, requiring effective integration of visual and linguistic information. In this work, we propose SegVLM, a vision-language model that incorporates architectural improvements to enhance segmentation accuracy and cross-modal alignment. The model integrates squeeze-and-excitation (SE) blocks for dynamic feature recalibration, deformable convolutions for geometric adaptability, and residual connections for deep feature learning. We also introduce a novel referring-aware fusion (RAF) loss that balances region-level alignment, boundary precision, and class imbalance. Extensive experiments and ablation studies demonstrate that each component contributes to consistent performance improvements. SegVLM also shows strong generalization across diverse datasets and referring expression scenarios.

指代分割视觉语言模型图像分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。