arXiv:2605.08557cs.CVcs.AI2026-05

用几何匹配的连续迁移提升少样本视觉模型适应能力

MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

论文配图:MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching
图 1 · 摘自论文原文
  • 将特征迁移建模为混合曲率流形上的连续过程
  • 在7个基准上多数设置表现最优,尤其对Transformer和细粒度数据集优势明显
  • 无需修改主干网络,适合快速部署于各类预训练模型

参数高效适配预训练视觉模型通常通过线性探测、提示词、低秩更新或轻量残差模块实现。这些方法常将适配视为对冻结特征的离散欧氏扰动,未显式建模任务引起的特征位移几何结构。我们提出MC-RFM,一种基于混合曲率黎曼流匹配的少样本适配框架,用于冻结视觉主干网络。核心思想是将适配特征表示在组合流形上:超曲面部分捕捉层次敏感的语义结构,欧氏部分保留局部判别性视觉变化。适配被形式化为从冻结特征到支持集原型的任务条件连续传输,通过流匹配目标训练,并耦合混合原型-线性分类器。该方法轻量、主干无关,完全在缓存的冻结特征上运行。在七个视觉识别基准、五种冻结主干及1/4/16样本设置下,MC-RFM在多数情形中表现最佳,尤其在Transformer主干和细粒度数据集上提升显著。消融实验表明,混合曲率头、任务条件、自适应分支门控、原型收缩与判别性监督均贡献性能。结果表明,少样本适配不仅取决于更新哪些参数,更依赖于如何在匹配下游任务结构的几何空间中移动表示。

原文摘要 · Abstract (English)

Parameter-efficient adaptation of pretrained vision models is commonly performed through linear probes, prompts, low-rank updates, or lightweight residual modules. While effective, these methods usually treat adaptation as a discrete Euclidean perturbation of frozen representations, without explicitly modeling the geometry of the task-induced feature displacement. We propose \textsc{MC-RFM}, a mixed-curvature Riemannian flow-matching framework for few-shot adaptation of frozen visual backbones. The key idea is to represent adapted features on a product manifold combining a hyperbolic factor, which captures hierarchy-sensitive semantic structure, and a Euclidean factor, which preserves locally discriminative visual variation. Adaptation is formulated as a task-conditioned continuous transport from frozen features to support-set prototypes, trained with a flow-matching objective and coupled to a hybrid prototype-linear classifier. The method is lightweight, backbone-agnostic, and operates entirely on cached frozen features. Across seven visual recognition benchmarks, five frozen backbones, and 1/4/16-shot regimes, \textsc{MC-RFM} is the best-performing method in a majority of evaluated settings, with the strongest gains on Transformer backbones and fine-grained datasets. Ablations show that the mixed-curvature head, task conditioning, adaptive branch gating, prototype shrinkage, and discriminative supervision each contribute to performance. These results suggest that few-shot adaptation benefits not only from deciding which parameters to update, but also from modeling how representations should move through a geometry matched to the structure of the downstream task.

少样本学习几何建模流匹配视觉适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。