通过分离模态融合提示,提升细粒度组合零样本学习性能
Separated Inter/Intra-Modal Fusion Prompts for Compositional Zero-Shot Learning
- 设计跨/内模态融合提示生成器,分离处理属性与对象关系
- 在CZSL任务中显著提升细粒度语义识别准确率,优于现有方法
- 适合研究视觉语言模型、零样本学习的科研人员使用
组合零样本学习(CZSL)旨在通过已知与未知概念的组合,在训练中识别细微语义差异或状态与物体的结合。现有方法要么专注于提示配置,要么利用提示微调预训练视觉-语言模型,但在精确识别细微语义差异或正确组合状态与对象方面仍存在挑战。为同时解决上述问题,本文提出一种新方法:通过引入多样化的提示学习,结合跨/内模态融合合成器,在涉及细微语义差异和多个对象的场景理解中,有效提升属性识别性能,构建更高效、精准的CZSL技术。
原文摘要 · Abstract (English)
Compositional Zero-Shot Learning (CZSL) aims to recognize subtle differences in meaning or the combination of states and objects through the use of known and unknown concepts during training. Existing methods either focused on prompt configuration or on using prompts to tune the pre-trained Vision-Language model. However, these methods faced challenges in accurately identifying subtle differences in meaning or combining states with objects. To jointly eradicate the above issues and construct an efficient and effective CZSL technique, we suggest a method to improve attribute recognition performance by utilizing diverse Prompt Learning with an Inter/Intra-Modality Fusion Synthesizer in scene understanding involving subtle semantic differences and multiple objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。