arXiv:2606.08231cs.CV2026-06ACL综述

探索多模态大模型推理时动态扩容的高效方法

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning

论文配图:Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning
图 1 · 摘自论文原文
  • 按采样、反馈、搜索三类策略统一分类多模态推理增强方法
  • 梳理典型任务与基准测试,揭示当前性能提升路径
  • 适合关注多模态推理优化的研究者和工程师参考

测试时缩放(Test-time Scaling, TTS)已成为通过动态分配推理阶段计算资源以提升模型性能的关键研究方向。近期进展已将这一范式拓展至多模态基础模型(Multimodal Foundation Models, MFMs),释放其在多模态推理与生成方面的潜力。尽管发展迅速,该领域仍缺乏系统性综述与统一理论框架来厘清演进脉络。为此,本文首次全面回顾了面向MFMs的TTS研究,提出一个统一的分类框架,将现有方法归纳为基于采样的、基于反馈的和基于搜索的三类策略。进一步总结了代表性应用及常用评估基准,用于衡量多模态TTS在生成与推理任务中的能力。最后,讨论了开放挑战并展望未来研究方向,为该快速发展的领域提供系统性路线图。

原文摘要 · Abstract (English)

Test-time Scaling (TTS) has emerged as a pivotal research direction for enhancing model performance by dynamically allocating computational resources during inference. Recent advancements have adapted this paradigm to Multimodal Foundation Models (MFMs), unlocking their potential in multimodal reasoning and generation. Despite rapid progress, the field lacks a systematic survey and unified theoretical framework to delineate the developmental landscape of multimodal TTS. To bridge this gap, we present the first comprehensive review of TTS research for MFMs, proposing a unified taxonomic framework that categorizes existing methodologies into three distinct strategies: sampling-based, feedback-based, and search-based approaches. We further summarize representative applications and benchmarks commonly utilized to evaluate multimodal TTS capabilities in generation and reasoning tasks. Finally, this survey discusses open challenges and outlines future research directions, providing a systematic roadmap for subsequent studies in this rapidly evolving field.

多模态推理优化测试时缩放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。