arXiv:2607.01535cs.CV2026-07

提出隐式提示机制,让视觉通用模型仅用一次新任务数据就能快速适应。

Hidden-Shot: Towards One-Shot Task Generalization for Low-Level Vision Generalist Models

论文配图:Hidden-Shot: Towards One-Shot Task Generalization for Low-Level Vision Generalist Models
图 1 · 摘自论文原文
  • 通过提取隐式视觉任务信息,构建全局文本提示增强泛化能力。
  • 在3C4U和3C7U评估框架下,单次任务学习性能超越现有最优模型。
  • 无需修改原模型结构,适合快速部署到各类低层视觉任务中。

尽管低层视觉通用模型备受关注,其在零样本/少样本场景下的新任务泛化能力仍缺乏验证。核心挑战在于实现从未见任务的泛化,且需匹配量化评估标准。现有方法虽在提示工程方面取得进展,但未系统探索跨多种低层视觉任务的泛化差距。为此,我们提出 Hidden-Shot,一种隐式提示机制,用于探索视觉通用模型在低层任务中的适应能力。该方法提取隐式视觉任务信息,利用全局任务感知文本提示,并选择性融合隐式信息与任务内处理信息,以提升新任务的一次学习能力。整体设计以低成本方式直接注入,几乎不改变原模型架构。此外,我们引入数据驱动的评估框架 C/U,涵盖两种基础场景:3C4U(3个常规+4个非常规任务)用于重训练现有模型,3C7U(3个常规+7个非常规任务)用于从零训练,全面测试低层通用模型的泛化能力。在七个和十个数据集上的实验分别优于当前最先进模型,经由3C4U和3C7U框架验证。Hidden-Shot 在新任务上表现出色,同时保持对已有任务的一致性能。

原文摘要 · Abstract (English)

Despite the intense engagement surrounding low-level vision generalist models, their effectiveness in zero/few-shot scenarios beyond learned tasks remains unverified. The primary challenge of developing an ideal generalist lies in achieving the ability to generalize from new unseen tasks, which also can be assessed by matched quantitative criteria. Existing methods have made some progress in prompt engineering but have not systematically explored this gap across a wide range of low-level visual tasks. Stimulated by the problem, we propose Hidden-Shot, an implicit prompt mechanism aimed at exploring low-level task adaptation in a vision generalist model. Specifically, the method extracts implicit visual task-based information, utilizes a global task-aware textural prompt, and selectively merges implicit information with in-task processing information to enhance one-shot capabilities in new tasks. The overall design performs direct injection in a cost-effective manner, while minimally altering the architecture of the original generalist model. Additionally, we introduce a data-driven evaluation framework termed C/U assessment to cover two basic scenarios, 3C4U (3 conventional and 4 unconventional tasks) for retraining existing models and 3C7U (3 conventional and 7 unconventional tasks) for training from scratch, as a comprehensive assessment to systematically test the generalization ability of low-level generalist models. Experiments on seven and ten datasets outperform the state-of-the-art vision generalist model, respectively verified by 3C4U and 3C7U framework. Our presented Hidden-Shot approach demonstrates superior performance on one-shot new tasks while maintaining consistent performance on existing tasks.

低层视觉一 Shot泛化能力提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。