arXiv:2504.15199cs.CVcs.AI2025-04

MILS框架零样本图文生成虽好,但计算成本极高。

Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning

  • 用迭代式LLM-CLIP方法实现零样本图像描述
  • 多步精炼导致计算开销远超BLIP-2和GPT-4V
  • 揭示零样本模型的隐性资源代价,适合关注效率的研究者

MILS(多模态迭代大语言模型求解器)是一种新提出的框架,宣称通过基于LLM-CLIP的迭代方法实现‘无需训练的视觉与听觉理解’,在零样本图像描述任务中表现良好。然而,我们发现其成功背后隐藏着巨大的计算开销,源于其复杂的多步精炼过程。相比之下,BLIP-2和GPT-4V等模型通过单次传递即可达成相近性能。我们推测,MILS的迭代机制带来的显著资源消耗可能削弱其实际应用价值,挑战了‘零样本无代价’的主流叙事。本文首次系统揭示并量化了MILS在输出质量与计算成本之间的权衡关系,为构建更高效的多模态模型提供关键洞见。

原文摘要 · Abstract (English)

MILS (Multimodal Iterative LLM Solver) is a recently published framework that claims "LLMs can see and hear without any training" by leveraging an iterative, LLM-CLIP based approach for zero-shot image captioning. While this MILS approach demonstrates good performance, our investigation reveals that this success comes at a hidden, substantial computational cost due to its expensive multi-step refinement process. In contrast, alternative models such as BLIP-2 and GPT-4V achieve competitive results through a streamlined, single-pass approach. We hypothesize that the significant overhead inherent in MILS's iterative process may undermine its practical benefits, thereby challenging the narrative that zero-shot performance can be attained without incurring heavy resource demands. This work is the first to expose and quantify the trade-offs between output quality and computational cost in MILS, providing critical insights for the design of more efficient multimodal models.

零样本多模态计算效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。