arXiv:2502.14786cs.CVcs.AI2025-02被引 1.2k

SigLIP 2提升多语言视觉语言理解,支持定位与密集预测。

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

  • 融合多种训练技巧的统一训练配方
  • 在零样本分类、检索等任务上全面超越原版
  • 支持多分辨率与原生长宽比,适合多场景应用

我们提出SigLIP 2,一个基于原始SigLIP成功经验的新一代多语言视觉-语言编码器家族。在第二代中,我们将原有的图像-文本训练目标扩展为统一训练方案,整合了基于标题的预训练、自监督损失(自蒸馏、掩码预测)和在线数据清洗等独立开发技术。通过这些改进,所有规模的SigLIP 2模型在核心能力上均优于其前身,包括零样本分类、图文检索以及用于视觉-语言模型的视觉表征迁移性能。此外,新训练方案显著提升了定位与密集预测任务表现。我们还训练了支持多分辨率且保持输入原始长宽比的变体,并使用更丰富的数据混合搭配去偏技术,大幅增强多语言理解能力与公平性。为平衡推理成本与性能,我们发布了四个尺寸的模型检查点:ViT-B(86M)、L(303M)、So400m(400M)和g(1B)。

原文摘要 · Abstract (English)

We introduce SigLIP 2, a family of new multilingual vision-language encoders that build on the success of the original SigLIP. In this second iteration, we extend the original image-text training objective with several prior, independently developed techniques into a unified recipe -- this includes captioning-based pretraining, self-supervised losses (self-distillation, masked prediction) and online data curation. With these changes, SigLIP 2 models outperform their SigLIP counterparts at all model scales in core capabilities, including zero-shot classification, image-text retrieval, and transfer performance when extracting visual representations for Vision-Language Models (VLMs). Furthermore, the new training recipe leads to significant improvements on localization and dense prediction tasks. We also train variants which support multiple resolutions and preserve the input's native aspect ratio. Finally, we train on a more diverse data-mixture that includes de-biasing techniques, leading to much better multilingual understanding and improved fairness. To allow users to trade off inference cost with performance, we release model checkpoints at four sizes: ViT-B (86M), L (303M), So400m (400M), and g (1B).

多语言视觉语言图像定位模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。