arXiv:2510.09201cs.LGcs.AI2025-10被引 4

让多模态输入参与提示优化,提升大模型表现

Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs

  • 提出联合优化文本与图像/视频等多模态提示的新框架
  • 在图像、视频、分子等场景中超越纯文本优化方法
  • 适合希望挖掘多模态大模型潜力的研究者

大语言模型(LLMs)已取得显著成功,其多模态扩展(MLLMs)进一步拓展了对图像、视频及其他非文本模态的能力。然而,尽管模态扩展已实现,提示优化方法仍局限于文本,限制了MLLM的全部潜力。为此,我们提出多模态提示优化新问题,将提示优化扩展至文本与非文本提示的组合空间。为此,我们设计了多模态提示优化器(MPO),一个统一框架,通过保持对齐的更新方式联合优化多模态提示,并基于贝叶斯策略利用早期评估结果作为先验指导候选提示选择。在涵盖图像、视频乃至分子等多种模态的广泛实验中,MPO显著优于主流纯文本优化方法,证明多模态提示优化是释放MLLM潜力的关键步骤。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable success, and their multimodal expansions (MLLMs) further unlock capabilities spanning images, videos, and other modalities beyond text. However, despite this shift, prompt optimization approaches, designed to reduce the burden of manual prompt crafting while maximizing performance, remain confined to text, ultimately limiting the full potential of MLLMs. Motivated by this gap, we introduce the new problem of multimodal prompt optimization, which expands the prior definition of prompt optimization to the multimodal space defined by the pairs of textual and non-textual prompts. To tackle this problem, we then propose the Multimodal Prompt Optimizer (MPO), a unified framework that not only performs the joint optimization of multimodal prompts through alignment-preserving updates but also guides the selection process of candidate prompts by leveraging earlier evaluations as priors in a Bayesian-based selection strategy. Through extensive experiments across diverse modalities that go beyond text, such as images, videos, and even molecules, we demonstrate that MPO outperforms leading text-only optimization methods, establishing multimodal prompt optimization as a crucial step to realizing the potential of MLLMs.

多模态提示优化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。