arXiv:2503.12999cs.CVcs.AI2025-03被引 2

用树结构生成可控的合成数据,提升视觉语言模型个性化能力。

Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs

  • 将概念建模为树结构,生成不同难度和多样性的正负样本。
  • 在多个基准上显著提升个性化VLM性能,解决正样本稀缺与负样本质量差问题。
  • 适合需要高效定制视觉语言模型的研究者与开发者使用。

视觉语言模型(VLM)在多模态任务中表现优异,近年来个性化能力受到关注。现有方法依赖用户提供的正负样本进行微调,但正样本稀缺、负样本质量低成为瓶颈。我们系统研究了正负样本数量与多样性(易/难样本)对个性化效果的影响。基于分析,提出概念即树(Concept-as-Tree, CaT)框架,将概念表示为树形结构,可生成具有不同难度与多样性的正负样本,并支持多概念扩展。结合精心设计的数据过滤策略,确保生成数据质量,构建强大合成数据流水线。在多种VLM个性化基线上的实验表明,配备该过滤器的CaT显著增强模型在个性化基准上的表现。据我们所知,这是首个面向VLM个性化的可控合成数据管道。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated exceptional performance in various multi-modal tasks. Recently, there has been an increasing interest in improving the personalization capabilities of VLMs. To better integrate user-provided concepts into VLMs, many methods use positive and negative samples to fine-tune these models. However, the scarcity of user-provided positive samples and the low quality of retrieved negative samples pose challenges for existing techniques. To reveal the relationship between sample and model performance, we systematically investigate the amount and diversity impact of positive and negative samples (easy and hard) on VLM personalization tasks. Based on the detailed analysis, we introduce Concept-as-Tree (CaT), which represents a concept as a tree structure, thereby enabling the data generation of positive and negative samples with varying difficulty and diversity, and can be easily extended to multi-concept scenarios. With a well-designed data filtering strategy, our CaT framework can ensure the quality of generated data, constituting a powerful pipeline. We perform thorough experiments with various VLM personalization baselines to assess the effectiveness of the pipeline, alleviating the lack of positive samples and the low quality of negative samples. Our results demonstrate that CaT equipped with the proposed data filter significantly enhances the capabilities of VLMs across personalization benchmarks. To the best of our knowledge, this work is the first controllable synthetic data pipeline for VLM personalization.

视觉语言模型个性化合成数据可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。