arXiv:2606.28194cs.LG2026-06

构建视觉逻辑推理数据集,精准定位模型判断失误点

COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives

论文配图:COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives
图 1 · 摘自论文原文
  • 按逻辑强度分正例、近边界、远边界负例,精细诊断模型表现
  • 模型对明显错误样本识别好,但对接近边界的模糊样本易出错
  • 适合研究可解释性、逻辑推理与少样本学习的学者

尽管可解释模型如概念瓶颈模型(CBMs)和程序合成方法能验证决策过程,其评估通常局限于简单任务,对真实图像中的复杂推理仍缺乏探索。我们提出COCOLogic-V2,一个面向真实图像的视觉归纳推理对象中心数据集,覆盖广泛的一阶逻辑子集。通过将样本分为正例、近边界(NB)和远边界(FB)负例,该数据集支持对模型可问责性的细粒度诊断。实验表明,模型能较好区分正例与FB负例,但在NB负例上表现不佳;感知噪声和大规模规则搜索空间在少样本设置下构成额外挑战。结果表明,视觉归纳推理仍是开放难题,而COCOLogic-V2为推进相关方法提供了坚实基础。

原文摘要 · Abstract (English)

While interpretable models such as concept bottleneck models (CBMs) and program synthesis methods enable verification of model decisions, their evaluation is typically limited to simple tasks, leaving complex reasoning on real-world images largely unexplored. We introduce COCOLogic-V2, an object-centric dataset for visual inductive reasoning on real-world images covering a broad subset of first-order logic. By categorizing samples into positive variants, near-boundary (NB), and far-from-boundary (FB) negatives, COCOLogic-V2 enables fine-grained diagnosis of model accountability. Our evaluations show that models tend to separate positive and FB samples well but fail on NB samples, while perceptual noise and large rule-induced search spaces pose additional challenges in few-shot settings. Together, these results highlight that visual inductive reasoning remains an open challenge and COCOLogic-V2 provides a concrete foundation for advancing methods in this direction.

视觉推理可解释性逻辑一致性少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。