arXiv:2509.16011cs.CV2025-09被引 1

用多个语义原型解决视觉持续学习中的歧义与多样性问题

Towards Robust Visual Continual Learning with Multi-Prototype Supervision

  • 用轻量级语言模型生成多语境原型,替代单一语义目标
  • 在多个基准上显著提升持续学习性能与鲁棒性
  • 适合需要应对类别歧义和视觉多样性的视觉持续学习场景

语言引导的监督通过使用预训练语言模型(PLM)中冻结的语义目标,已成为视觉持续学习(CL)的一种有前景范式。然而,依赖单一目标会带来两个关键限制:1)语义模糊性,即多义类别名称导致视觉表征冲突;2)类内视觉多样性,即单个原型无法捕捉类别内部丰富的外观变化。为此,我们提出 MuproCL,一种新框架,将单一目标替换为多个上下文感知的原型。具体地,我们采用轻量级 LLM 代理进行类别消歧和视觉模态扩展,生成稳健的语义原型集。通过 LogSumExp 聚合机制,视觉模型可自适应地与最相关的原型对齐。在多种 CL 基线上的广泛实验表明,MuproCL 持续提升性能与鲁棒性,为语言引导的持续学习提供了更有效的路径。

原文摘要 · Abstract (English)

Language-guided supervision, which utilizes a frozen semantic target from a Pretrained Language Model (PLM), has emerged as a promising paradigm for visual Continual Learning (CL). However, relying on a single target introduces two critical limitations: 1) semantic ambiguity, where a polysemous category name results in conflicting visual representations, and 2) intra-class visual diversity, where a single prototype fails to capture the rich variety of visual appearances within a class. To this end, we propose MuproCL, a novel framework that replaces the single target with multiple, context-aware prototypes. Specifically, we employ a lightweight LLM agent to perform category disambiguation and visual-modal expansion to generate a robust set of semantic prototypes. A LogSumExp aggregation mechanism allows the vision model to adaptively align with the most relevant prototype for a given image. Extensive experiments across various CL baselines demonstrate that MuproCL consistently enhances performance and robustness, establishing a more effective path for language-guided continual learning.

视觉持续学习多原型语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。