arXiv:2502.19973cs.CVcs.AI2025-02被引 2

测试大模型在复杂场景中整合多模态信息的推理能力,发现现有模型表现不佳。

Can Large Language Models Unveil the Mysteries? An Exploration of Their Ability to Unlock Information in Complex Scenarios

  • 设计双基准测试,评估多模态信息组合推理能力
  • SOTA模型在复杂推理任务上准确率仅33.04%和7.38%
  • 新方法提升22.17%和9.40%,适合多模态推理研究者

人类在复杂场景中结合多重感知输入并进行组合推理是一种高级认知功能。随着多模态大语言模型的发展,现有基准测试多聚焦于跨图像的视觉理解,却常忽略对多源感知信息组合推理的需求。为探索先进模型在复杂场景中融合多模态输入进行组合推理的能力,我们引入两个基准:Clue-Visual Question Answering(CVQA),包含三种任务类型以评估视觉理解和综合;Clue of Password-Visual Question Answering(CPVQA),包含两种任务类型,聚焦于对视觉数据的精确解读与应用。针对这些基准,我们提出三种即插即用的方法:利用模型输入进行推理、通过最小间隔解码结合随机生成增强推理、检索语义相关的视觉信息以实现有效数据融合。联合结果表明,当前模型在组合推理基准上的表现较差,即使是最先进的闭源模型在CVQA上也仅达33.04%准确率,而在CPVQA上降至7.38%。值得注意的是,我们的方法使模型在组合推理任务上的性能显著提升,在CVQA上比SOTA模型提高22.17%,在CPVQA上提高9.40%,证明了其在复杂场景中融合多模态输入进行组合推理的有效性。代码将公开。

原文摘要 · Abstract (English)

Combining multiple perceptual inputs and performing combinatorial reasoning in complex scenarios is a sophisticated cognitive function in humans. With advancements in multi-modal large language models, recent benchmarks tend to evaluate visual understanding across multiple images. However, they often overlook the necessity of combinatorial reasoning across multiple perceptual information. To explore the ability of advanced models to integrate multiple perceptual inputs for combinatorial reasoning in complex scenarios, we introduce two benchmarks: Clue-Visual Question Answering (CVQA), with three task types to assess visual comprehension and synthesis, and Clue of Password-Visual Question Answering (CPVQA), with two task types focused on accurate interpretation and application of visual data. For our benchmarks, we present three plug-and-play approaches: utilizing model input for reasoning, enhancing reasoning through minimum margin decoding with randomness generation, and retrieving semantically relevant visual information for effective data integration. The combined results reveal current models' poor performance on combinatorial reasoning benchmarks, even the state-of-the-art (SOTA) closed-source model achieves only 33.04% accuracy on CVQA, and drops to 7.38% on CPVQA. Notably, our approach improves the performance of models on combinatorial reasoning, with a 22.17% boost on CVQA and 9.40% on CPVQA over the SOTA closed-source model, demonstrating its effectiveness in enhancing combinatorial reasoning with multiple perceptual inputs in complex scenarios. The code will be publicly available.

多模态推理大模型评测组合推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。