arXiv:2512.16202cs.CVcs.AI2025-12CVPR被引 3

让AI动态学习新类别,比现有方法更准更透明。

Open Ad-hoc Categorization with Contextualized Feature Learning

  • 用可学习的上下文标记引导冻结的CLIP模型进行语义扩展。
  • 在Stanford Mood数据集上达到87.4%的新类别准确率,超越基线超50%。
  • 生成可解释的注意力图,适合需要透明决策的场景应用。

自适应视觉场景分类对AI代理应对变化任务至关重要。与植物、动物等固定类别不同,临时类别是为特定目标动态创建的。本文研究开放型临时分类:给定少量标注样本和大量未标注数据,目标是发现潜在上下文,并通过语义扩展和视觉聚类来拓展临时类别。基于临时类别与常见类别依赖相似感知机制的洞察,我们提出OAK模型——在冻结的CLIP输入端引入一组可学习的上下文标记,联合优化CLIP的图文对齐目标与GCD的视觉聚类目标。在Stanford和Clevr-4数据集上,OAK在多个分类任务中实现最优性能,包括在Stanford Mood上达到87.4%的新型类别准确率,较CLIP和GCD提升超过50%。此外,OAK生成可解释的显著性图,聚焦于动作中的手、情绪中的脸、位置中的背景,提升透明度与可信度,支持自适应且泛化性强的分类。

原文摘要 · Abstract (English)

Adaptive categorization of visual scenes is essential for AI agents to handle changing tasks. Unlike fixed common categories for plants or animals, ad-hoc categories are created dynamically to serve specific goals. We study open ad-hoc categorization: Given a few labeled exemplars and abundant unlabeled data, the goal is to discover the underlying context and to expand ad-hoc categories through semantic extension and visual clustering around it. Building on the insight that ad-hoc and common categories rely on similar perceptual mechanisms, we propose OAK, a simple model that introduces a small set of learnable context tokens at the input of a frozen CLIP and optimizes with both CLIP's image-text alignment objective and GCD's visual clustering objective. On Stanford and Clevr-4 datasets, OAK achieves state-of-the-art in accuracy and concept discovery across multiple categorizations, including 87.4% novel accuracy on Stanford Mood, surpassing CLIP and GCD by over 50%. Moreover, OAK produces interpretable saliency maps, focusing on hands for Action, faces for Mood, and backgrounds for Location, promoting transparency and trust while enabling adaptive and generalizable categorization.

图像分类动态学习可解释性CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。