arXiv:2503.06847cs.CV2025-03被引 2

用多属性文档监督提升零样本图像分类,过滤非视觉噪声

MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification

  • 用大模型自动清理文档中的非视觉描述,保留视觉相关属性
  • 在三个基准上比现有最优方法分别提升7.2%和8.2%准确率
  • 适合做零样本图像识别且需要可解释预测的研究者

零样本学习(ZSL)旨在通过共享辅助信息实现从已见类别到未见类别的知识迁移。近期研究发现,百科全书文档提供有用辅助信息。但现有方法将混杂视觉与非视觉描述的嘈杂文档与图像区域对齐,仅依赖隐式学习,无法可靠过滤非视觉噪声,导致非视觉词错误对齐图像区域,损害知识迁移。本文提出多属性文档监督框架MADS,从文档收集与模型学习双阶段去除噪声。借助大语言模型,设计新型提示算法,自动剔除非视觉描述并多属性丰富低描述文档。MADS通过信息解耦与语义交互,在局部与全局层次提取多视角可迁移知识。此外,引入模型无关聚焦损失,显式增强训练中对视觉判别信息的关注,无需额外参数即可提升现有方法性能。在计算成本相当情况下,MADS在三个基于文档的ZSL与通用零样本学习(GZSL)基准上平均提升7.2%和8.2%。同时,提供多属性视角下的可解释预测。

原文摘要 · Abstract (English)

Zero-shot learning (ZSL) aims to train a model on seen classes and recognize unseen classes by knowledge transfer through shared auxiliary information. Recent studies reveal that documents from encyclopedias provide helpful auxiliary information. However, existing methods align noisy documents, entangled in visual and non-visual descriptions, with image regions, yet solely depend on implicit learning. These models fail to filter non-visual noise reliably and incorrectly align non-visual words to image regions, which is harmful to knowledge transfer. In this work, we propose a novel multi-attribute document supervision framework to remove noises at both document collection and model learning stages. With the help of large language models, we introduce a novel prompt algorithm that automatically removes non-visual descriptions and enriches less-described documents in multiple attribute views. Our proposed model, MADS, extracts multi-view transferable knowledge with information decoupling and semantic interactions for semantic alignment at local and global levels. Besides, we introduce a model-agnostic focus loss to explicitly enhance attention to visually discriminative information during training, also improving existing methods without additional parameters. With comparable computation costs, MADS consistently outperforms the SOTA by 7.2% and 8.2% on average in three benchmarks for document-based ZSL and GZSL settings, respectively. Moreover, we qualitatively offer interpretable predictions from multiple attribute views.

零样本学习文档监督多属性可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。