用随机标签训练让语言模型适配视觉任务,效果超预期。
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
- 用随机标签做桥接训练,无需人工标注即可对齐语言与视觉参数。
- 部分层的预训练参数在视觉任务中仍具强基础能力,无需全量微调。
- 为跨模态迁移提供新路径,适合想复用语言模型做视觉任务的研究者。
语言预训练模型与视觉预训练模型的异常参数比例差异显著,导致跨模态(语言与视觉)比跨领域适应更具挑战性。以往研究多聚焦于跨领域迁移,认为语言预训练模型因参数空间差异不适于下游视觉任务。与此相反,我们发现通过引入桥接训练阶段作为模态适配学习器,可有效对齐大语言模型(LLM)参数与视觉任务。具体提出一种无需人工标注的简单而高效的方法——随机标签桥接训练,助力LLM参数适配视觉基础任务。此外,研究发现部分桥接训练更优:某些LLM层具有强基础特性,在不针对视觉任务微调时依然有效。这一意外发现为直接利用语言预训练参数于视觉模型开辟新途径,并凸显部分桥接训练在跨模态适配中的实用潜力。
原文摘要 · Abstract (English)
The ratio of outlier parameters in language pre-training models and vision pre-training models differs significantly, making cross-modality (language and vision) inherently more challenging than cross-domain adaptation. As a result, many prior studies have focused on cross-domain transfer rather than attempting to bridge language and vision modalities, assuming that language pre-trained models are unsuitable for downstream visual tasks due to disparate parameter spaces. Contrary to this assumption, we show that adding a bridge training stage as a modality adaptation learner can effectively align Large Language Model (LLM) parameters with vision tasks. Specifically, we propose a simple yet powerful solution random label bridge training that requires no manual labeling and helps LLM parameters adapt to vision foundation tasks. Moreover, our findings reveal that partial bridge training is often advantageous, as certain layers in LLMs exhibit strong foundational properties that remain beneficial even without fine-tuning for visual tasks. This surprising discovery opens up new avenues for leveraging language pre-trained parameters directly within vision models and highlights the potential of partial bridge training as a practical pathway to cross-modality adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。