通过异常对比学习预测图像组合关系中的异常项
Predictive Reasoning with Augmented Anomaly Contrastive Learning for Compositional Visual Relations
- 设计增强型异常对比学习,提取通用判别特征
- 在3个图像中预测第4个,准确率显著超越现有模型
- 适合需要逻辑推理与模式识别的视觉任务研究者
尽管简单类比的视觉推理已受到广泛关注,但因复杂度更高,组合视觉关系(CVR)仍较少被研究。为解决CVR任务,本文提出预测性推理与增强异常对比学习(PR-A²CL),即在三个遵循相同组合规律的图像中识别出异常项。针对组合规则繁多的挑战,设计了增强异常对比学习,通过最大化正常样本间的相似性、最小化正常与异常样本间的相似性,提炼出判别性强且泛化能力高的特征。更重要的是,引入“预测-验证”范式,利用一系列预测异常推理模块(PARBs)迭代地基于三张图像的特征预测第四张图像的特征,并在后续验证阶段逐步定位由底层规则导致的具体差异。在SVRT、CVR和MC²R数据集上的实验结果表明,PR-A²CL显著优于当前最先进的推理模型。
原文摘要 · Abstract (English)
While visual reasoning for simple analogies has received significant attention, compositional visual relations (CVR) remain relatively unexplored due to their greater complexity. To solve CVR tasks, we propose Predictive Reasoning with Augmented Anomaly Contrastive Learning (PR-A$^2$CL), \ie, to identify an outlier image given three other images that follow the same compositional rules. To address the challenge of modelling abundant compositional rules, an Augmented Anomaly Contrastive Learning is designed to distil discriminative and generalizable features by maximizing similarity among normal instances while minimizing similarity between normal and anomalous outliers. More importantly, a predict-and-verify paradigm is introduced for rule-based reasoning, in which a series of Predictive Anomaly Reasoning Blocks (PARBs) iteratively leverage features from three out of the four images to predict those of the remaining one. Throughout the subsequent verification stage, the PARBs progressively pinpoint the specific discrepancies attributable to the underlying rules. Experimental results on SVRT, CVR and MC$^2$R datasets show that PR-A$^2$CL significantly outperforms state-of-the-art reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。