提升视觉语言模型推理能力,用新方法实现更优测试时计算。
Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models

- 基于预测熵选择最可信输出,动态优化多模型集成决策。
- 在6个数据集上超越单模型与传统投票法,小模型可增强大模型性能。
- 揭示多样性缺失是传统方法效果有限的关键原因,适合模型集成研究者。
测试时计算(TTC)策略已成为提升大型语言模型推理能力的轻量级方法,但其在视觉语言模型(VLMs)中的应用与价值尚未充分探索。我们系统性地研究了七种VLM和六个基准上的TTC,重点分析基于特征的评分与多数投票方法。发现特征启发式方法失效,单一模型下投票仅带来微弱增益。理论证明,该局限源于预测缺乏多样性:当输出高度相关时,投票无法显著改进。相比之下,多模型集成具有更高多样性,但标准多数投票未考虑模型能力差异。为此,我们提出基于熵的TTC(ETTC),根据预测熵选择最自信的输出。在单模型情况下退化为多数投票,而在多模型中利用置信度差异优先强模型。我们证明在温和假设下ETTC优于多数投票,并实证显示其始终超越投票法及最优单模型。关键发现:小模型可协同增强大模型,实现标准策略无法达成的集成收益。
原文摘要 · Abstract (English)
Test-time compute (TTC) strategies have emerged as a lightweight approach to boost reasoning in large language models (LLMs). However, their application and benefits for vision-language models (VLMs) remain underexplored. We present a systematic study of TTC across seven VLMs and six benchmarks, specifically analyzing feature-based scoring and majority voting methods. We find that feature heuristics fail and voting yields only modest gains in single-model settings. We theoretically show that this limitation stems from a lack of prediction diversity: when outputs are highly correlated, voting provides little benefit. In contrast, multi-model ensembles offer richer diversity, yet standard majority voting fails to account for varying model capabilities. To address this, we propose Entropy-based TTC (ETTC), which selects the most confident prediction based on predictive entropy. Our method reduces to majority voting in the single-model case, but in model ensembles, it leverages confidence disparities to prioritize stronger models. We prove that ETTC outperforms majority voting under mild assumptions and empirically demonstrate that it consistently surpasses both voting and the best individual model. Crucially, our results show that smaller models can synergistically enhance larger ones, unlocking ensembling gains not achievable with standard strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。