通过双向纠错提升组合零样本学习的推理能力
Progressive Reasoning with Primitive Correction for Compositional Zero-Shot Learning

- 分步推理+双向校正,解决属性与物体预测的依赖关系
- 在三个基准上达到最优性能,显著降低错误传播
- 适合关注组合泛化与逻辑推理的视觉语言研究者
组合零样本学习(CZSL)旨在将已知的属性与物体作为基本单元,识别之前未见过的属性-物体组合。以往方法或独立预测属性与物体,忽略其强上下文依赖;或采用单向条件建模(如物体引导属性预测),易导致错误传播。本文提出PRPC框架——基于原始单元修正的渐进式推理,通过分步推理显式建模属性与物体间的双向依赖,并在每一步进行原始单元的相互校正以抑制早期预测误差。我们将CZSL建模为结构化的问答式思维链过程,约束多模态大模型按预定义语义步骤生成中间决策。为进一步增强中间推理的可靠性和逻辑一致性,引入基于GRPO的强化学习后训练,提供与渐进推理过程对齐的步骤级奖励。在三个CZSL基准上的大量实验表明,PRPC实现当前最优性能,验证了渐进推理与双向校正在鲁棒组合泛化中的有效性。
原文摘要 · Abstract (English)
Compositional Zero-Shot Learning (CZSL) aims to combine known attributes and objects as primitives for recognizing previously unseen attribute-object pairs. Prior works either predict attributes and objects independently, missing their strong contextual dependency, or use unidirectional conditional modeling (e.g., object-guided attribute prediction), which is prone to error propagation. We propose PRPC, a Progressive Reasoning framework with Primitive Correction, which explicitly models the bidirectional dependency between attributes and objects via step-wise inference. PRPC performs mutual correction of primitives to suppress prediction errors in earlier steps. Specifically, we formulate CZSL as structured, Q&A-style Chain-of-Thought reasoning process and constrain the MLLM to follow predefined semantic steps to generate intermediate decisions. To further enhance the reliability and logical consistency of intermediate reasoning, we introduce reinforcement learning post-training with a GRPO-based objective, providing step-level rewards aligned with the progressive inference procedure. Extensive experiments on three CZSL benchmarks demonstrate that PRPC achieves state-of-the-art performance, validating the effectiveness of progressive reasoning and bidirectional correction for robust compositional generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。