提出复杂视觉查询,测试模型抽象推理能力。
Symbolic and Abstractive Reasoning with Complex Visual Queries

- 设计复杂视觉查询(CVQ)数据类型,融合一阶逻辑构造多样问题
- 构建含14类查询的合成数据集,支持大规模训练与评估
- 两阶段训练框架提升模型跨任务泛化能力,适合研究推理机制者
当前多模态大模型在理解抽象视觉内容方面仍具挑战。本文提出一种新型抽象数据类型——复杂视觉查询(CVQ),用于探测符号与抽象推理能力,这是人类神经符号推理的关键但未充分探索维度。我们从数据、范式与探索三方面开展系统研究:提出基于大规模多模态知识图谱的可扩展合成管道,通过一阶逻辑运算符系统组合生成涵盖14种不同类型的多样化数据集;进一步设计两阶段训练框架,逐步赋予多模态大模型强健的视觉推理能力。通过多维度实验,严格评估模型在CVQ推理性能、跨任务与跨场景泛化能力等方面表现。本工作为推进多模态大模型的推理边界提供了新视角与路径。
原文摘要 · Abstract (English)
Understanding and reasoning over abstract visual content remains a challenge for current multi-modal large language models (MLLMs). In this paper, we explore a novel abstract data type termed complex visual query (CVQ), designed to probe symbolic and abstractive reasoning, which is a critical yet underexplored dimension of human-like neuro-symbolic reasoning for MLLMs. We present a comprehensive investigation from three perspectives: \textbf{Data $\times$ Paradigm $\times$ Exploration}. Specifically, we propose a scalable pipeline for synthesizing CVQs grounded in large-scale multi-modal knowledge graphs, generating a diverse dataset encompassing 14 distinct query types via systematic combinations of first-order logic operators. We further introduce a two-stage training framework that progressively equips MLLMs with robust visual reasoning capabilities. We conduct extensive experiments to rigorously evaluate MLLMs across multiple dimensions, including reasoning performance on CVQs, as well as cross-task and cross-scenario generalization. We believe our work opens new perspectives and avenues for advancing the reasoning frontiers of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。