用视觉基础模型重新定义泛化分割,让模型自适应现实场景变化。
Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
- 以视觉基础模型为枢纽,动态优化特征适应下游任务。
- 在基准数据集上显著优于现有方法,尤其在泛化分割任务中表现突出。
- 适合研究视觉模型泛化与自适应推理的学者和工程师。
本文首次提出集合枢轴学习(Set Pivot Learning, SPL)的概念,推动基于视觉基础模型(VFMs)的领域泛化范式革新。传统领域泛化假设目标域不可见,但大规模多样化数据训练的视觉基础模型使该假设过时。SPL通过动态适应与模型中心调优,实现从静态对齐到任务驱动特征优化的转变,支持模型随真实场景持续演化。具体包含两个关键特性:(i) 动态适应,由固定域对齐转向灵活的任务导向特征优化;(ii) VFM中心调优,利用预训练知识作为枢纽,强化特定任务表征同时保持跨域鲁棒性。在此基础上,提出动态提示微调方法,结合类感知动态提示器与提示引导特征聚焦器,提升视觉基础模型在目标场景中的性能。大量实验验证了方法有效性,尤其在泛化分割任务中显著超越当前最优方法。
原文摘要 · Abstract (English)
In this paper, we introduce, for the first time, the concept of Set Pivot Learning, a paradigm shift that redefines domain generalization (DG) based on Vision Foundation Models (VFMs). Traditional DG assumes that the target domain is inaccessible during training, but the emergence of VFMs, trained on vast and diverse data, renders this assumption unclear and obsolete. Traditional DG assumes that the target domain is inaccessible during training, but the emergence of VFMs, which are trained on vast and diverse datasets, renders this assumption unclear and obsolete. To address this challenge, we propose Set Pivot Learning (SPL), a new definition of domain migration task based on VFMs, which is more suitable for current research and application requirements. Unlike conventional DG methods, SPL prioritizes adaptive refinement over rigid domain transfer, ensuring continuous alignment with evolving real-world conditions. Specifically, SPL features two key attributes: (i) Dynamic adaptation, transitioning from static domain alignment to flexible, task-driven feature optimization, enabling models to evolve with downstream scenarios; (ii) VFM-centric tuning, leveraging pretrained knowledge as a pivot to hone task-specific representations while preserving cross-domain robustness. Building on SPL, we propose a Dynamic Prompt Fine-Tuning method, which combines a Dynamic Class-aware Prompter with a Prompt-guided Feature Focuser, to elevate VFM performance in targeted scenarios. Extensive experiments on benchmark datasets show the effectiveness of our method, highlighting its superiority over state-of-the-art methods, particularly in generalized segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。