arXiv:2507.00868cs.CV2025-07ICCV被引 3

让视觉模型在测试时灵活组合医疗任务,无需重训。

Is Visual in-Context Learning for Compositional Medical Tasks within Reach?

  • 用合成任务生成器训练模型适应连续任务序列。
  • 在多个分割数据集上实现复杂医疗任务组合,无需重新训练。
  • 适合需要灵活调整视觉分析流程的医疗场景研究者。

本文探索视觉上下文学习在组合式医疗任务中的潜力,旨在让单一模型在测试时通过无须重训的方式处理多种任务并动态适应新任务。不同于以往针对单个任务的方法,本文关注模型对任务序列的适应能力,目标是通过统一模型完成涉及多个中间步骤的复杂任务,使用户可在测试时灵活定义视觉处理流程。为此,首先分析了视觉上下文学习架构的特性与局限性,特别是代码本(codebooks)的作用;随后提出一种基于合成组合任务生成引擎的新训练方法,该引擎从任意分割数据集构建任务序列,从而支持视觉上下文学习模型在组合任务上的训练。此外,还研究了不同基于掩码的训练目标,以深入理解如何更有效地训练模型解决复杂组合任务。研究不仅为多模态医疗任务序列提供了重要见解,也揭示了仍需克服的关键挑战。

原文摘要 · Abstract (English)

In this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks during test time without re-training. Unlike previous approaches, our focus is on training in-context learners to adapt to sequences of tasks, rather than individual tasks. Our goal is to solve complex tasks that involve multiple intermediate steps using a single model, allowing users to define entire vision pipelines flexibly at test time. To achieve this, we first examine the properties and limitations of visual in-context learning architectures, with a particular focus on the role of codebooks. We then introduce a novel method for training in-context learners using a synthetic compositional task generation engine. This engine bootstraps task sequences from arbitrary segmentation datasets, enabling the training of visual in-context learners for compositional tasks. Additionally, we investigate different masking-based training objectives to gather insights into how to train models better for solving complex, compositional tasks. Our exploration not only provides important insights especially for multi-modal medical task sequences but also highlights challenges that need to be addressed.

视觉推理医疗图像上下文学习组合任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。