arXiv:2410.12816cs.CVcs.CL2024-10NeurIPS被引 15

从因果视角解决视觉语言模型适配中的数据错位问题。

Rethinking Misalignment in Vision-Language Model Adaptation from a Causal Perspective

  • 构建因果模型分析任务无关知识干扰
  • 提出CDC方法解耦语义并降低预测不确定性
  • 适用于需高精度分类的下游视觉任务

基础视觉语言模型如CLIP在下游任务中表现出色,但在适配特定任务时存在两级错位问题:任务错位与数据错位。软提示微调缓解了任务错位,但数据错位仍难解决。本文重新审视CLIP的预训练与适配过程,建立结构化因果模型,发现尽管期望准确捕捉任务相关特征,但任务无关知识会干扰预测结果,阻碍图像与类别间真实关系建模。由于任务无关知识不可观测,我们采用前门调整,提出因果引导的语义解耦与分类(CDC)方法。具体地,对下游任务数据中的语义进行解耦,并基于每种语义分别分类;同时利用Dempster-Shafer证据理论评估各语义预测的不确定性。在多种设置下的实验一致验证了CDC的有效性。

原文摘要 · Abstract (English)

Foundational Vision-Language models such as CLIP have exhibited impressive generalization in downstream tasks. However, CLIP suffers from a two-level misalignment issue, i.e., task misalignment and data misalignment, when adapting to specific tasks. Soft prompt tuning has mitigated the task misalignment, yet the data misalignment remains a challenge. To analyze the impacts of the data misalignment, we revisit the pre-training and adaptation processes of CLIP and develop a structural causal model. We discover that while we expect to capture task-relevant information for downstream tasks accurately, the task-irrelevant knowledge impacts the prediction results and hampers the modeling of the true relationships between the images and the predicted classes. As task-irrelevant knowledge is unobservable, we leverage the front-door adjustment and propose Causality-Guided Semantic Decoupling and Classification (CDC) to mitigate the interference of task-irrelevant knowledge. Specifically, we decouple semantics contained in the data of downstream tasks and perform classification based on each semantic. Furthermore, we employ the Dempster-Shafer evidence theory to evaluate the uncertainty of each prediction generated by diverse semantics. Experiments conducted in multiple different settings have consistently demonstrated the effectiveness of CDC.

视觉语言模型因果推理语义解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。