arXiv:2503.17940cs.CV2025-03CVPR被引 27

用费舍尔信息引导微调,提升视觉大模型在跨域分割中的泛化能力

FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation

  • 基于领域相关费舍尔信息矩阵,动态识别敏感参数并选择性更新
  • 在多个跨域分割数据集上达到最优性能,显著优于传统微调方法
  • 适合需要强泛化能力的视觉模型迁移任务,尤其关注领域适应场景

视觉基础模型(VFMs)凭借大规模预训练展现出优异的泛化能力,但在面向领域泛化语义分割(DGSS)进行微调时,如何保持其泛化性仍具挑战。现有方法或仅选择性微调参数,或冻结模型仅更新适配器,均可能未能充分挖掘VFMs在DGSS任务中的潜力。我们观察到,由任务与分布差异引发的领域敏感参数会损害泛化性能。为此,提出FisherTune,一种由领域相关费舍尔信息矩阵(DR-FIM)引导的鲁棒微调方法。DR-FIM衡量参数在不同任务与领域下的敏感度,实现有选择性的参数更新,从而保留泛化能力并增强对DGSS的适应性。FisherTune引入变分推断以稳定DR-FIM估计,将参数视为高斯分布变量,并利用预训练先验。大量实验表明,FisherTune在跨域分割任务中表现更优,同时维持了良好的泛化性,显著超越选择性参数与适配器基方法。

原文摘要 · Abstract (English)

Vision Foundation Models (VFMs) excel in generalization due to large-scale pretraining, but fine-tuning them for Domain Generalized Semantic Segmentation (DGSS) while maintaining this ability remains challenging. Existing approaches either selectively fine-tune parameters or freeze the VFMs and update only the adapters, both of which may underutilize the VFMs' full potential in DGSS tasks. We observe that domain-sensitive parameters in VFMs, arising from task and distribution differences, can hinder generalization. To address this, we propose \textbf{FisherTune}, a robust fine-tuning method guided by the Domain-Related Fisher Information Matrix (DR-FIM). DR-FIM measures parameter sensitivity across tasks and domains, enabling selective updates that preserve generalization and enhance DGSS adaptability. FisherTune incorporates variational inference to stabilize DR-FIM estimation, treating parameters as Gaussian-distributed variables and leveraging pre-trained priors. Extensive experiments show that FisherTune achieves superior cross-domain segmentation while maintaining generalization, outperforming selective-parameter and adapter-based methods.

视觉模型领域泛化微调分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。