arXiv:2512.11109cs.LG2025-12被引 4

测试时扩展能提升视觉语言模型推理能力,但效果因模型和任务而异。

Limits and Gains of Test-Time Scaling in Vision-Language Reasoning

  • 通过测试时增加计算量,采用结构化推理与迭代自修正
  • 闭源模型普遍受益,开源模型仅外部验证有效且迭代可能退化
  • 多步推理任务提升明显,感知类任务增益有限,需适配策略

测试时扩展(TTS)作为提升大语言模型推理能力的有力范式,通过在推理阶段分配额外计算实现性能增强,但其在视觉语言模型(VLMs)中的应用仍缺乏系统研究。本文对开源与闭源VLM在不同基准上的推理方法进行了系统性实证分析。结果表明:闭源模型持续受益于结构化推理与迭代自修正;而开源VLM表现不一致——外部验证带来最可靠提升,迭代精炼常导致性能下降。此外,TTS效果具有数据集依赖性,在多步推理任务中显著提升,但在以感知为核心的基准上增益有限。研究揭示了TTS并非通用解法,必须结合模型能力和任务特性进行定制,推动未来自适应TTS策略与多模态奖励模型的发展。

原文摘要 · Abstract (English)

Test-time scaling (TTS) has emerged as a powerful paradigm for improving the reasoning ability of Large Language Models (LLMs) by allocating additional computation at inference, yet its application to multimodal systems such as Vision-Language Models (VLMs) remains underexplored. In this work, we present a systematic empirical study of inference time reasoning methods applied across both open-source and closed-source VLMs on different benchmarks. Our results reveal that while closed-source models consistently benefit from structured reasoning and iterative Self-Refinement, open-source VLMs show inconsistent behavior: external verification provides the most reliable gains, whereas iterative refinement often degrades performance. We further find that the effectiveness of TTS is dataset-dependent, yielding clear improvements on multi-step reasoning tasks but offering only limited gains on perception-focused benchmarks. These findings demonstrate that TTS is not a universal solution and must be tailored to both model capabilities and task characteristics, motivating future work on adaptive TTS strategies and multimodal reward models.

视觉语言测试时扩展推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。