arXiv:2511.18121cs.CVcs.AI2025-11被引 2

提出层次化视觉内涵理解框架,让大模型像人一样逐步推理。

VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging

  • 构建从感知到抽象的分层推理机制,实现线索到结论的可追溯推断。
  • 在多级推理中发现性能随层级提升持续下降,揭示模型瓶颈。
  • 通过蒙特卡洛树搜索生成数据,强化底层能力可提升整体表现。

尽管多模态大语言模型在基准测试中表现优异,但其处理方式与人类整合视觉信息的能力存在差异。人类能自然地连接细节与高层概念,而模型常将二者割裂。现有评估协议往往将低层感知与高层推理分离,忽视其语义与因果关联,导致结果不可诊断且难以定位性能瓶颈。本文提出VCU-Bridge框架,模拟人类层次化的视觉内涵理解:从基础感知经由语义桥接,到抽象内涵的多层级推理,并建立从具体线索到抽象结论的显式证据-推断链。基于此框架,我们构建了HVCU-Bench,一个具有层级诊断能力的基准。全面实验显示,随着推理层级提升,性能持续下降。我们进一步设计了基于蒙特卡洛树搜索(MCTS)的指令微调数据生成管道,结果显示增强低层能力可显著提升高层表现。有趣的是,该改进不仅在HVCU-Bench上有效,还带来通用基准平均+2.53%的提升,尤其在MMStar上达到+7.26%,验证了层次化思维模式的重要性及其对模型能力的有效提升。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visual information. Unlike humans who naturally bridge details and high-level concepts, models tend to treat these elements in isolation. Prevailing evaluation protocols often decouple low-level perception from high-level reasoning, overlooking their semantic and causal dependencies, which yields non-diagnostic results and obscures performance bottlenecks. We present VCU-Bridge, a framework that operationalizes a human-like hierarchy of visual connotation understanding: multi-level reasoning that advances from foundational perception through semantic bridging to abstract connotation, with an explicit evidence-to-inference trace from concrete cues to abstract conclusions. Building on this framework, we construct HVCU-Bench, a benchmark for hierarchical visual connotation understanding with explicit, level-wise diagnostics. Comprehensive experiments demonstrate a consistent decline in performance as reasoning progresses to higher levels. We further develop a data generation pipeline for instruction tuning guided by Monte Carlo Tree Search (MCTS) and show that strengthening low-level capabilities yields measurable gains at higher levels. Interestingly, it not only improves on HVCU-Bench but also brings benefits on general benchmarks (average +2.53%), especially with substantial gains on MMStar (+7.26%), demonstrating the significance of the hierarchical thinking pattern and its effectiveness in enhancing MLLM capabilities. The project page is at https://vcu-bridge.github.io .

多模态推理框架层次化评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。