发现抽象视觉推理中感知细节是大模型主要瓶颈,提出新数据生成方法提升性能。
VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
- 构建细粒度抽象视觉推理基准VisuRiddles,覆盖五维能力与两类高阶推理。
- 设计自动合成器生成带精细描述的谜题,显著提升模型在新任务上的准确率。
- 适合关注多模态模型视觉理解、可解释性研究的研究者使用。
近年来,多模态大语言模型(MLLMs)在多项推理任务中取得显著进展,但抽象视觉推理(AVR)仍是关键挑战,主要源于对抽象图形的感知能力不足。为此,我们探究当前MLLMs的瓶颈,并合成训练数据以增强其抽象视觉感知能力。首先,提出VisuRiddles基准,包含精心设计的任务,用于评估模型在五个核心维度和两类高阶推理上的推理能力。其次,引入感知谜题生成器(PRS),一种自动化框架,可生成带有细粒度感知描述的谜题。该方法不仅产生有价值的抽象图形训练数据,还提供中间推理阶段的监督信号,从而提升训练效率与模型可解释性。在VisuRiddles上的大量实验表明,细粒度视觉感知是主要瓶颈,而我们的合成框架显著提升了现有MLLMs在这些挑战性任务中的表现。代码与数据集将开源于https://github.com/yh-hust/VisuRiddles。
原文摘要 · Abstract (English)
Recent strides in multimodal large language models (MLLMs) have significantly advanced their performance in many reasoning tasks. However, Abstract Visual Reasoning (AVR) remains a critical challenge, primarily due to limitations in perceiving abstract graphics. To tackle this issue, we investigate the bottlenecks in current MLLMs and synthesize training data to improve their abstract visual perception. First, we propose VisuRiddles, a benchmark for AVR, featuring tasks meticulously constructed to assess models' reasoning capacities across five core dimensions and two high-level reasoning categories. Second, we introduce the Perceptual Riddle Synthesizer (PRS), an automated framework for generating riddles with fine-grained perceptual descriptions. PRS not only generates valuable training data for abstract graphics but also provides fine-grained perceptual description, crucially allowing for supervision over intermediate reasoning stages and thereby improving both training efficacy and model interpretability. Our extensive experimental results on VisuRiddles empirically validate that fine-grained visual perception is the principal bottleneck and our synthesis framework markedly enhances the performance of contemporary MLLMs on these challenging tasks. Our code and dataset will be released at https://github.com/yh-hust/VisuRiddles
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。