无需调参和少量数据,快速适配私有领域高效多模态模型。
Boosting Private Domain Understanding of Efficient MLLMs: A Tuning-free, Adaptive, Universal Prompt Optimization Framework
- 基于强化搜索生成提示优化策略树,建立先验知识。
- 通过自反思机制动态优化提示,适应私有数据分布。
- 零参数微调,适合资源受限设备的隐私保护场景。
高效多模态大模型(EMLLMs)相比传统多模态大模型(MLLMs)具有更小的模型规模与更低的计算开销,常部署于资源受限设备。然而,由于数据隐私限制,现有开源EMLLMs在预训练阶段通常无法访问私有领域数据,难以直接应用于特定业务场景。为解决此问题,本文提出一种无需调参、自适应、通用的提示优化框架 extit{ extbf{ hiswork{}}},包含两个阶段:1)基于强化搜索策略生成提示优化策略树,获取优化先验;2)基于先验初始化提示,并通过自反思机制进一步搜索与精炼提示。该方法无需参数微调,仅需少量数据即可快速适应私有数据分布。大量实验表明,相较于基线方法, hiswork{}在多项任务中显著提升了效率与性能。
原文摘要 · Abstract (English)
Efficient multimodal large language models (EMLLMs), in contrast to multimodal large language models (MLLMs), reduce model size and computational costs and are often deployed on resource-constrained devices. However, due to data privacy concerns, existing open-source EMLLMs rarely have access to private domain-specific data during the pre-training process, making them difficult to directly apply in device-specific domains, such as certain business scenarios. To address this weakness, this paper focuses on the efficient adaptation of EMLLMs to private domains, specifically in two areas: 1) how to reduce data requirements, and 2) how to avoid parameter fine-tuning. Specifically, we propose a tun\textbf{\underline{I}}ng-free, a\textbf{\underline{D}}aptiv\textbf{\underline{E}}, univers\textbf{\underline{AL}} \textbf{\underline{Prompt}} Optimization Framework, abbreviated as \textit{\textbf{\ourmethod{}}} which consists of two stages: 1) Predefined Prompt, based on the reinforcement searching strategy, generate a prompt optimization strategy tree to acquire optimization priors; 2) Prompt Reflection initializes the prompt based on optimization priors, followed by self-reflection to further search and refine the prompt. By doing so, \ourmethod{} elegantly generates the ``ideal prompts'' for processing private domain-specific data. Note that our method requires no parameter fine-tuning and only a small amount of data to quickly adapt to the data distribution of private data. Extensive experiments across multiple tasks demonstrate that our proposed \ourmethod{} significantly improves both efficiency and performance compared to baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。