首个面向多模态的自动提示优化框架,解决视觉令牌膨胀与过程监督缺失问题。
UniAPO: Unified Multimodal Automated Prompt Optimization
- 采用类似EM算法的解耦优化流程,分离反馈建模与提示优化。
- 在文本、图像、视频任务上均实现稳定提升,跨模态表现优异。
- 适合需要高效提示工程的多模态模型开发者与研究者。
提示工程是释放大语言模型潜力的关键。为自动化并增强该过程,自动提示优化(APO)已被提出,主要在纯文本场景中表现出色。然而,将现有APO方法扩展至多模态任务(如视频-语言生成)面临两大挑战:(i) 视觉令牌膨胀,长视觉序列限制上下文容量,导致反馈信号不足;(ii) 缺乏过程级监督,现有方法仅关注结果级监督,忽视中间步骤,限制了提示优化效果。我们提出UniAPO:统一多模态自动提示优化框架,首个专为多模态APO设计的系统。UniAPO采用类EM优化流程,解耦反馈建模与提示精炼,使优化更稳定且目标导向。为应对上述挑战,引入长短时记忆机制:历史反馈缓解上下文限制,历史提示提供优化方向。UniAPO在文本、图像、视频基准上均实现一致提升,建立了高效可迁移的统一提示优化框架。
原文摘要 · Abstract (English)
Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, demonstrating effectiveness primarily in text-only input scenarios. However, extending existing APO methods to multimodal tasks, such as video-language generation introduces two core challenges: (i) visual token inflation, where long visual token sequences restrict context capacity and result in insufficient feedback signals; (ii) a lack of process-level supervision, as existing methods focus on outcome-level supervision and overlook intermediate supervision, limiting prompt optimization. We present UniAPO: Unified Multimodal Automated Prompt Optimization, the first framework tailored for multimodal APO. UniAPO adopts an EM-inspired optimization process that decouples feedback modeling and prompt refinement, making the optimization more stable and goal-driven. To further address the aforementioned challenges, we introduce a short-long term memory mechanism: historical feedback mitigates context limitations, while historical prompts provide directional guidance for effective prompt optimization. UniAPO achieves consistent gains across text, image, and video benchmarks, establishing a unified framework for efficient and transferable prompt optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。