提出统一多模态上下文学习框架,提升少样本任务适应能力
UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy
- 构建六级能力分类体系,系统化定义示范作用
- 创建76万条8样本示范数据集,显著提升跨任务性能
- 轻量级模块增强稳定性,适合多模态少样本研究者使用
上下文学习(ICL)可在不更新参数的情况下快速适配新任务,但对示范样本选择和格式高度敏感。在涵盖理解与生成的统一多模态模型中,这种敏感性因跨模态干扰和认知负荷差异而加剧,导致性能非单调且任务依赖性强。为此,我们提出六级能力导向分类体系,将示范功能从基础感知到高阶判断进行系统划分。基于该认知框架,构建UniICL-760K大规模语料库,包含15个子任务、8样本示范的76万条上下文学习样本,并配套设计UniICL-Bench进行严格控制评估。实验证明,数据驱动的构建是性能提升主因。此外,提出轻量级上下文自适应原型调制器(CAPM),作为插件模块进一步提升少样本稳定性。在UniICL-Bench上,该方法在多数理解类任务中优于更大参数的多模态大语言模型基线。数据与代码已开源。
原文摘要 · Abstract (English)
In-context learning (ICL) enables fast task adaptation from demonstrations without per-task parameter updates but remains highly sensitive to example selection and formatting. In unified multimodal models spanning understanding and generation, this sensitivity is exacerbated by cross-modal interference and varying cognitive demands. Consequently, in-context learning efficacy is often non-monotonic and highly task-dependent. To diagnose these behaviors, we introduce a six-level Capability-Oriented Taxonomy that categorizes the functional role of demonstrations from basic perception to high-order discernment. Guided by this cognitive framework, we construct UniICL-760K, a large-scale corpus featuring curated 8-shot in-context learning episodes across 15 subtasks, alongside UniICL-Bench for rigorous, controlled evaluation. We show that this data-driven assembly is the primary source of our gains. As a complementary, lightweight stabilizer, we additionally propose the Context-Adaptive Prototype Modulator, a plug-and-play module that further improves few-shot stability. Evaluations on UniICL-Bench show that our approach yields highly competitive unified results, outperforming larger-parameter multimodal large language model baselines on most understanding in-context learning tasks. Data and code are available at https://github.com/xuyicheng-zju/UniICL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。