通过特征矩阵提升视觉语言模型在通用任务上的表现
Enhancing Target-unspecific Tasks through a Features Matrix
- 构建特征矩阵捕捉深层通用知识,防止过拟合
- 在跨数据集、跨域等通用任务上达到顶尖性能
- 可无缝集成到现有框架,适合通用视觉任务研究者
大型视觉语言模型的提示学习虽显著提升了特定任务表现,但在通用或非目标相关任务上仍显不足,原因在于训练过拟合导致模型遗忘通用知识。为此,我们提出一种新型特征矩阵(FM)方法,旨在增强模型在通用任务中的能力。该方法从深层细粒度角度提取并利用通用知识,构建特征矩阵,有效保留关键通用信息,降低过拟合风险。实验表明:1)FM可作为通用灵活模块兼容现有框架;2)在基础到新类别泛化、域泛化和跨数据集泛化等目标无关任务中表现卓越,达到当前最优水平。
原文摘要 · Abstract (English)
Recent developments in prompt learning of large Vision-Language Models (VLMs) have significantly improved performance in target-specific tasks. However, these prompting methods often struggle to tackle the target-unspecific or generalizable tasks effectively. It may be attributed to the fact that overfitting training causes the model to forget its general knowledge. The general knowledge has a strong promotion on target-unspecific tasks. To alleviate this issue, we propose a novel Features Matrix (FM) approach designed to enhance these models on target-unspecific tasks. Our method extracts and leverages general knowledge, shaping a Features Matrix (FM). Specifically, the FM captures the semantics of diverse inputs from a deep and fine perspective, preserving essential general knowledge, which mitigates the risk of overfitting. Representative evaluations demonstrate that: 1) the FM is compatible with existing frameworks as a generic and flexible module, and 2) the FM significantly showcases its effectiveness in enhancing target-unspecific tasks (base-to-novel generalization, domain generalization, and cross-dataset generalization), achieving state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。