arXiv:2412.12940cs.CL2024-12中稿 · ed被引 4

仅用文本训练就能提升视觉模型细粒度识别能力。

Improving Fine-grained Visual Understanding in VLMs through Text-Only Training

  • 用纯文本数据训练视觉语言模型,替代图像-文本配对数据。
  • 在物种分类和文化视觉理解任务上达到与传统方法相当的准确率。
  • 显著降低计算成本,适合资源受限场景使用。

视觉语言模型(VLMs)已成为连接视觉与语言理解的强大工具。然而,传统VLM训练方法常受限于图像-文本配对数据的高收集与训练成本。近期研究指出,语言理解在VLM性能中起关键作用,暗示纯文本训练可能可行。本文探究通过纯文本训练提升VLM细粒度视觉理解的可行性。受人类通过丰富文本描述建立视觉概念的启发,我们假设VLM可借助文本表征增强视觉识别能力。在细粒度物种分类和文化视觉理解两个不同领域进行综合实验,结果表明:纯文本训练可达到与图像-文本训练相当的性能,同时大幅降低计算开销。这为提升VLM能力提供了一条更高效、低成本的新路径,尤其适用于资源受限环境。

原文摘要 · Abstract (English)

Visual-Language Models (VLMs) have become a powerful tool for bridging the gap between visual and linguistic understanding. However, the conventional learning approaches for VLMs often suffer from limitations, such as the high resource requirements of collecting and training image-text paired data. Recent research has suggested that language understanding plays a crucial role in the performance of VLMs, potentially indicating that text-only training could be a viable approach. In this work, we investigate the feasibility of enhancing fine-grained visual understanding in VLMs through text-only training. Inspired by how humans develop visual concept understanding, where rich textual descriptions can guide visual recognition, we hypothesize that VLMs can also benefit from leveraging text-based representations to improve their visual recognition abilities. We conduct comprehensive experiments on two distinct domains: fine-grained species classification and cultural visual understanding tasks. Our findings demonstrate that text-only training can be comparable to conventional image-text training while significantly reducing computational costs. This suggests a more efficient and cost-effective pathway for advancing VLM capabilities, particularly valuable in resource-constrained environments.

视觉语言模型文本训练细粒度识别低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。