arXiv:2509.11840cs.CV2025-09ICCV被引 1

用生成式模型造描述,实现零样本分割的精准语义对齐

Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation

论文配图:Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
图 1 · 摘自论文原文
  • 用生成式视觉语言模型生成图像的合成描述,实现图文空间对齐
  • 在标准零样本开放词汇分割任务上性能超越现有方法,且更省数据
  • 适合研究视觉语言对齐、零样本分割或追求高效标注的学者

生成式视觉语言模型(VLMs)具备强大的高层图像理解能力,但缺乏视觉与语言模态之间的空间密集对齐,我们的研究发现如此。与生成式VLMs的发展并行,另一方向的研究聚焦于视觉语言对齐的表征学习,旨在实现分割等密集任务的零样本推理。本文通过将图像与由VLM生成的合成描述进行密集对齐,融合两条技术路径。合成描述成本低、可扩展、易生成,是密集对齐方法获取高层语义理解的理想来源。实验证明,该方法在标准零样本开放词汇分割基准测试中优于先前工作,同时具有更高的数据效率。

原文摘要 · Abstract (English)

Generative vision-language models (VLMs) exhibit strong high-level image understanding but lack spatially dense alignment between vision and language modalities, as our findings indicate. Orthogonal to advancements in generative VLMs, another line of research has focused on representation learning for vision-language alignment, targeting zero-shot inference for dense tasks like segmentation. In this work, we bridge these two directions by densely aligning images with synthetic descriptions generated by VLMs. Synthetic captions are inexpensive, scalable, and easy to generate, making them an excellent source of high-level semantic understanding for dense alignment methods. Empirically, our approach outperforms prior work on standard zero-shot open-vocabulary segmentation benchmarks/datasets, while also being more data-efficient.

零样本分割视觉语言对齐生成式模型开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。