无需标注数据,让视觉模型像人一样分步推理。
CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT
- 分层合成思维链,从局部到整体构建逻辑结构。
- 在多层级语义不一致测试中达83.33%准确率。
- 适合需要可解释、强泛化视觉推理的场景。
视觉语言模型(VLMs)虽显著提升了图像与文本对齐能力,但仍难以实现类人视觉推理。主要瓶颈在于依赖表面相关性,缺乏逻辑连贯的结构化表征,导致高层语义结构缺失与非因果关系理解,阻碍组合式与可验证推理。为引入人类认知机制,我们提出CoTZero——一种无标注范式,包含两个部分:(i) 双阶段数据合成方法;(ii) 认知对齐训练策略。底层阶段受神经认知理论启发,提取原子视觉要素并逐步组合成多样化、结构化的问答-推理形式;顶层阶段利用粗粒度全局结构引导局部细节与因果关系解读。在认知对齐训练中,基于合成的思维链数据,引入认知一致可验证奖励(CCVR),通过强化微调提供分步反馈,提升推理连贯性与事实正确性。实验表明,CoTZero在包含词汇扰动负样本的多层级语义不一致基准上,在域内与域外设置下均取得83.33%的F1分数。消融实验证明各组件协同提升可解释性与人类对齐推理能力。
原文摘要 · Abstract (English)
Recent advances in vision-language models (VLMs) have markedly improved image-text alignment, yet they still fall short of human-like visual reasoning. A key limitation is that many VLMs rely on surface correlations rather than building logically coherent structured representations, which often leads to missed higher-level semantic structure and non-causal relational understanding, hindering compositional and verifiable reasoning. To address these limitations by introducing human models into the reasoning process, we propose CoTZero, an annotation-free paradigm with two components: (i) a dual-stage data synthesis approach and (ii) a cognition-aligned training method. In the first component, we draw inspiration from neurocognitive accounts of compositional productivity and global-to-local analysis. In the bottom-up stage, CoTZero extracts atomic visual primitives and incrementally composes them into diverse, structured question-reasoning forms. In the top-down stage, it enforces hierarchical reasoning by using coarse global structure to guide the interpretation of local details and causal relations. In the cognition-aligned training component, built on the synthesized CoT data, we introduce Cognitively Coherent Verifiable Rewards (CCVR) in Reinforcement Fine-Tuning (RFT) to further strengthen VLMs' hierarchical reasoning and generalization, providing stepwise feedback on reasoning coherence and factual correctness. Experiments show that CoTZero achieves an F1 score of 83.33 percent on our multi-level semantic inconsistency benchmark with lexical-perturbation negatives, across both in-domain and out-of-domain settings. Ablations confirm that each component contributes to more interpretable and human-aligned visual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。