分析视觉语言模型任务迁移规律,发现哪些任务能互相促进或干扰。
Understanding Task Transfer in Vision-Language Models
- 提出归一化指标PGF,量化一个任务微调对其他任务的影响。
- 在13个感知任务上构建迁移图谱,揭示任务间的正负迁移关系。
- 发现任务分组与行为模式,指导高效数据选择和训练策略。
视觉语言模型(VLMs)在多模态基准上表现良好,但在深度估计、物体计数等视觉感知任务上仍落后于人类和专用模型。对某一任务进行微调可能不可预测地影响其他任务的表现,使得任务特定微调面临挑战。本文通过系统研究任务可迁移性,探讨在某一感知任务上微调VLM如何影响其在其他任务上的零样本性能。我们引入完美差距因子(Perfection Gap Factor, PGF),一种归一化度量,用于衡量任务迁移导致的性能变化。利用PGF计算任务迁移性,该指标同时捕捉迁移的广度与强度。基于三个开源权重VLM在13个感知任务上的评估,我们构建了任务迁移图谱,揭示了此前未被观察到的任务间关系。分析发现正向与负向迁移模式,识别出相互影响的任务组,并根据迁移行为将任务归纳为不同‘人格’类型。此外,展示了PGF如何指导数据选择以实现更高效的训练。这些发现凸显了正向迁移的机会与负向干扰的风险,为推进VLM发展提供了可操作的指导。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like depth estimation or object counting. Finetuning on one task can unpredictably affect performance on others, making task-specific finetuning challenging. In this paper, we address this challenge through a systematic study of task transferability. We examine how finetuning a VLM on one perception task affects its zero-shot performance on others. We introduce Perfection Gap Factor (PGF), a normalized metric that measures change in performance as a result of task transfer. We utilize PGF to compute Task Transferability, which captures both the breadth and the magnitude of transfer induced by a source task. Using three open-weight VLMs evaluated across 13 perception tasks, we construct a task transfer graph that reveals previously unobserved relationships among perception tasks. Our analysis uncovers patterns of positive and negative transfer, identifies groups of tasks that mutually influence each other, organizes tasks into personas based on their transfer behavior and demonstrates how PGF can guide data selection for more efficient training. These findings highlight both opportunities for positive transfer and risks of negative interference, offering actionable guidance for advancing VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。