arXiv:2607.12786cs.CV2026-07中稿 · ACMMM2026

提升视觉语言模型跨图比较推理能力,解决细粒度属性定位难题。

CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

论文配图:CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models
图 1 · 摘自论文原文
  • 构建三元组训练数据集与结构化奖励机制,协同优化属性定位与判断一致性。
  • 在新基准上实现28.2分部分准确率提升,显著超越最强基线。
  • 适合关注多模态细粒度推理与跨图对比任务的研究者使用。

跨图像比较推理对视觉语言模型仍具挑战性,尤其当正确预测需细粒度属性定位和全局一致推理时。本文提出统一框架CoRe,包含:(i) CoRe-20K,一个从结构化视觉元数据自动构建的大规模三元组训练集,覆盖计数、深度、距离与空间关系;(ii) TriSR,一种结构化奖励框架,在GRPO优化下联合监督属性定位、判断对齐与三元组一致性;(iii) CoRe-Bench,首个专注细粒度跨图像比较推理的基准。实验表明,CoRe在CoRe-Bench上显著优于现有VLMs,同时在标准多模态基准上保持竞争力,部分准确率较最强基线提升28.2点。

原文摘要 · Abstract (English)

Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified framework for this problem. CoRe includes: (i) CoRe-20K, a large-scale triplet-based training set automatically constructed from structured visual metadata through a multi-expert collaborative pipeline, covering counting, depth, distance, and spatial relations; (ii) TriSR, a structured reward framework that jointly supervises attribute grounding, judgment alignment, and triplet consistency under GRPO optimization; and (iii) CoRe-Bench, the first benchmark dedicated to fine-grained cross-image comparative reasoning. Experiments show that CoRe substantially outperforms existing VLMs on CoRe-Bench while remaining competitive on standard multimodal benchmarks, achieving a 28.2-point gain in partial accuracy over the strongest baseline.

跨图推理视觉语言模型细粒度识别多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。