解决标签模糊下的零样本识别问题,提升模型鲁棒性。
Dynamic Visual-semantic Alignment for Zero-shot Learning with Ambiguous Labels

- 通过双向视觉语义对齐与注意力机制校准特征与原型。
- 基于互信息的对比优化增强语义一致性属性,提升判别力。
- 动态标签消歧迭代修正噪声,适合真实场景弱监督任务。
零样本学习(ZSL)旨在不依赖图像实例的情况下识别未见类别。然而,现有方法通常假设标签干净,忽略了现实中的标签噪声与模糊性,导致性能下降。为此,我们提出动态视觉-语义对齐(DVSA)框架,用于处理模糊标签的鲁棒性ZSL。DVSA采用双向视觉-语义对齐模块结合注意力机制,相互校准视觉特征与属性原型;同时,在属性层面基于互信息(MI)进行对比优化,强化具有判别性的语义一致属性。此外,动态标签消歧机制通过迭代修正噪声监督信号,保持语义一致性,缩小实例与标签间的差距,提升泛化能力。在标准基准上的大量实验表明,DVSA在模糊标注条件下实现更强性能。
原文摘要 · Abstract (English)
Zero-shot learning (ZSL) aims to recognize unseen classes without visual instances. However, existing methods usually assume clean labels, overlooking real-world label noise and ambiguity, which degrades performance. To bridge this gap, we propose the Dynamic Visual-semantic Alignment (DVSA), a robust ZSL framework for learning from ambiguous labels. DVSA uses a bidirectional visual-semantic alignment module with attention to mutually calibrate visual features and attribute prototypes, and a contrastive optimization grounded in Mutual Information (MI) at the attribute level to strengthen discriminative, semantically consistent attributes. In addition, a dynamic label disambiguation mechanism iteratively corrects noisy supervision while preserving semantic consistency, narrowing the instance-label gap, and improving generalization. Extensive experiments on standard benchmarks verify that DVSA achieves stronger performance under ambiguous supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。