arXiv:2510.12931cs.CVcs.CL2025-10中稿 · NeurIPS被引 1

无需标签数据,让视觉语言模型生成更准确的图像描述。

Unifying Vision-Language Latents for Zero-label Image Caption Enhancement

  • 通过对齐视觉与语言特征,实现无标签数据下的图像描述增强。
  • 在SmolVLM-Base和Qwen2-VL上提升描述准确性,细节更丰富。
  • 适合希望提升模型描述能力又无法获取标注数据的研究者。

视觉语言模型(VLMs)通过大规模图文预训练取得了显著性能,但其依赖标注图像数据集限制了可扩展性,大量未标注图像数据未被充分利用。为此,我们提出统一视觉语言表征对齐框架ViZer,实现图像描述任务中的零标签学习,为视觉语言任务的零标签适配提供实用起点。与依赖人工或合成标注数据的方法不同,ViZer在训练中主动对齐视觉与语言表示特征,使现有VLMs能在无需文本标签或全量重训练的前提下生成更优描述。定性评估显示,该方法优于传统指标(如CIDEr、BERTScore),后者常因参考描述缺失细节而惩罚合理新增内容。在SmolVLM-Base和Qwen2-VL上应用后,生成描述更具事实依据且更富描述性。

原文摘要 · Abstract (English)

Vision-language models (VLMs) achieve remarkable performance through large-scale image-text pretraining. However, their reliance on labeled image datasets limits scalability and leaves vast amounts of unlabeled image data underutilized. To address this, we propose Unified Vision-Language Alignment for Zero-Label Enhancement (ViZer), an enhancement training framework that enables zero-label learning in image captioning, providing a practical starting point for broader zero-label adaptation in vision-language tasks. Unlike prior approaches that rely on human or synthetically annotated datasets, ViZer actively aligns vision and language representation features during training, enabling existing VLMs to generate improved captions without requiring text labels or full retraining. We demonstrate ViZer's advantage in qualitative evaluation, as automated caption metrics such as CIDEr and BERTScore often penalize details that are absent in reference captions. Applying ViZer on SmolVLM-Base and Qwen2-VL, we observe consistent qualitative improvements, producing captions that are more grounded and descriptive than their baseline.

图像描述零样本学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。