提升图像视频交互分割精度,保持稳定提示能力
Towards Fine-grained Interactive Segmentation in Images and Videos
- 通过跨注意力增强局部上下文,优化全局特征
- 引入提示重定向模块,强化对细粒度目标的响应
- 多尺度级联结构生成高分辨率精确掩码,适合精细分割场景
近期的通用交互式分割模型(如SAM)虽具备强泛化能力,但在需要高精度掩码的场景中仍表现下降。现有方法在捕捉局部细节与保持稳定提示能力之间存在权衡,限制了基础分割模型的应用效果。为此,我们提出基于SAM2骨干网络的SAM2Refiner框架,可在图像和视频上生成细粒度分割掩码,同时保留其原有优势。具体地,设计定位增强模块,利用跨注意力机制融合局部上下文信息以增强全局特征,挖掘潜在细节模式并维持语义一致性;为强化对增强嵌入的提示能力,引入提示重定向模块,用空间对齐的提示特征更新嵌入表示;此外,通过多尺度级联结构融合编码器的分层特征,构建掩码精修模块以生成高分辨率精确掩码。大量实验表明,该方法在图像与视频任务上均优于现有最先进方法。
原文摘要 · Abstract (English)
The recent Segment Anything Models (SAMs) have emerged as foundational visual models for general interactive segmentation. Despite demonstrating robust generalization abilities, they still suffer performance degradations in scenarios demanding accurate masks. Existing methods for high-precision interactive segmentation face a trade-off between the ability to perceive intricate local details and maintaining stable prompting capability, which hinders the applicability and effectiveness of foundational segmentation models. To this end, we present an SAM2Refiner framework built upon the SAM2 backbone. This architecture allows SAM2 to generate fine-grained segmentation masks for both images and videos while preserving its inherent strengths. Specifically, we design a localization augment module, which incorporates local contextual cues to enhance global features via a cross-attention mechanism, thereby exploiting potential detailed patterns and maintaining semantic information. Moreover, to strengthen the prompting ability toward the enhanced object embedding, we introduce a prompt retargeting module to renew the embedding with spatially aligned prompt features. In addition, to obtain accurate high resolution segmentation masks, a mask refinement module is devised by employing a multi-scale cascaded structure to fuse mask features with hierarchical representations from the encoder. Extensive experiments demonstrate the effectiveness of our approach, revealing that the proposed method can produce highly precise masks for both images and videos, surpassing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。