用一张图就能让模型学会用户自定义的视觉任务,无需重训练。
Personalized Vision via Visual In-Context Learning
- 基于扩散模型构建四面板框架,单图示例即可推断变换规则。
- 在新任务上超越微调方法,支持开放式个性化需求。
- 适合需要快速适配新视觉任务的研究者与开发者。
当前视觉模型在大规模标注数据集上表现优异,但在个性化视觉任务(如用户自定义对象或新目标)上表现不佳。现有个性化方法依赖昂贵的微调或合成数据流程,灵活性差且仅限固定任务格式。视觉上下文学习(Visual ICL)提供了替代方案,但以往方法局限于狭窄、同域任务,无法泛化至开放式个性化。本文提出个人化上下文操作器(PICO),一个四面板框架,将扩散变压器改造为视觉上下文学习器。给定单张标注示例,PICO 能推断底层变换并应用于新输入,无需重新训练。为此,我们构建了紧凑而多样化的调优数据集 VisRel,证明任务多样性比数据规模更关键。此外,提出注意力引导的种子评分器,通过高效推理扩展提升可靠性。大量实验表明,PICO (i) 超越微调与合成数据基线,(ii) 灵活适应用户自定义的新任务,(iii) 在识别与生成任务间实现跨领域泛化。
原文摘要 · Abstract (English)
Modern vision models, trained on large-scale annotated datasets, excel at predefined tasks but struggle with personalized vision -- tasks defined at test time by users with customized objects or novel objectives. Existing personalization approaches rely on costly fine-tuning or synthetic data pipelines, which are inflexible and restricted to fixed task formats. Visual in-context learning (ICL) offers a promising alternative, yet prior methods confine to narrow, in-domain tasks and fail to generalize to open-ended personalization. We introduce Personalized In-Context Operator (PICO), a simple four-panel framework that repurposes diffusion transformers as visual in-context learners. Given a single annotated exemplar, PICO infers the underlying transformation and applies it to new inputs without retraining. To enable this, we construct VisRel, a compact yet diverse tuning dataset, showing that task diversity, rather than scale, drives robust generalization. We further propose an attention-guided seed scorer that improves reliability via efficient inference scaling. Extensive experiments demonstrate that PICO (i) surpasses fine-tuning and synthetic-data baselines, (ii) flexibly adapts to novel user-defined tasks, and (iii) generalizes across both recognition and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。