arXiv:2510.07632cs.AIcs.CL2025-10被引 5

通过测试时匹配提升多模态模型的组合推理能力

Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models

  • 提出测试时匹配算法,自迭代优化模型推理性能
  • 使SigLIP-B16超越GPT-4.1,在MMVP-VLM上创纪录
  • 无需外部监督,适用于多种模型与复杂数据集

前沿AI模型虽进展显著,但现有评估指标常低估其组合推理能力。本文提出群体匹配评分,更真实反映模型表现,并通过简单过拟合将新指标正确性转化为旧指标正确性。基于此,提出测试时匹配(TTM)算法,一种无需外部监督的自迭代优化方法。TTM使SigLIP-B16在MMVP-VLM上超越GPT-4.1,创下新纪录;同时在生成式多模态模型上也取得显著提升。该方法在16个不同数据集变体上均持续有效,尤其在WhatsUp等难题上实现最高85.7%的相对增益,大幅推进了组合推理能力边界。

原文摘要 · Abstract (English)

Frontier AI models have achieved remarkable progress, yet recent studies suggest they struggle with compositional reasoning, often performing at or below random chance on established benchmarks. We revisit this problem and show that widely used evaluation metrics systematically underestimate model capability. To correct this artifact, we introduce a group matching score that more faithfully evaluates model capability. Moreover, correctness under the new metric can be translated into correctness under existing metrics via a simple overfitting step. This adjustment enables SigLIP-B16 to surpass all previous results and GPT-4.1 to yield the first result surpassing estimated human performance on Winoground. Building on this insight, we propose Test-Time Matching (TTM), an iterative, self-improving algorithm that further bootstraps model performance without any external supervision. TTM delivers additional, non-trivial improvements: for example, TTM enables SigLIP-B16 to surpass GPT-4.1 on MMVP-VLM, establishing a new state of the art. TTM also extends beyond contrastive vision-language models, yielding clear gains on a generative multimodal model across benchmarks. Importantly, TTM remains broadly effective even on benchmarks without metric-induced effects or group structures, achieving relative gains up to 85.7% on challenging datasets such as WhatsUp. Across 16 dataset variants spanning diverse setups, our experiments demonstrate that TTM consistently improves model performance and advances the frontier of compositional reasoning.

多模态组合推理测试时优化自提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。