arXiv:2608.12724cs.LG2026-08

用无标签多模态数据提升少样本提示学习效果

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

论文配图:MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning
图 1 · 摘自论文原文
  • 构建多模态图,通过半监督传播筛选关键未标注样本
  • 在少标签场景下,相比基线模型显著提升性能
  • 适合资源受限但需高效适应新任务的研究者

基于多模态大模型的少样本提示学习(ICL)可在不更新参数的情况下完成任务适配,但其表现高度依赖演示样本的质量与覆盖范围。尽管无标签多模态数据丰富,如何有效利用仍不明确。本文提出MAG(MAnifold-Guided semi-supervised in-context demonstration selection),一种高效框架,利用无标签数据增强多模态ICL。MAG将演示选择建模为多模态图上的半监督传播问题,采用两阶段策略:(i) 相关性得分传播识别出少量高影响力未标注样本进行伪标签,降低MLLM推理开销;(ii) 结合多模态相关性最终选定演示样本。实验表明,文本表示更适用于相关性传播,而视觉与文本模态共同决定高质量演示选择。在八个多模态基准上,MAG在标签稀缺情况下持续优于强基线,在有限伪标签预算下实现显著提升。

原文摘要 · Abstract (English)

Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. We propose MAG (MAnifold-Guided semi-supervised in-context demonstra- tion selection), an efficient framework that leverages unlabeled data to improve multi-modal ICL. MAG formulates demonstration selection as a semi-supervised propagation problem on a multi-modal graph and adopts a two-stage strategy: (i) relevance score propagation identifies a compact set of high-impact unlabeled samples for pseudo-labeling, reducing MLLM inference cost; (ii) multi-modal relevance is used to select the final demonstrations. We show that textual represen- tations are more effective for relevance propagation, while both visual and textual modalities are crucial for high-quality demonstration selection. Experiments on eight multi-modal benchmarks demonstrate that MAG consistently outperforms strong baselines in label-scarce regimes, achieving significant gains with a limited pseudo-labeling budget.

提示学习多模态半监督少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。