arXiv:2509.17562cs.CV2025-09被引 8

用推理能力提升视觉模型感知,专攻遥感与医疗影像

Visual Instruction Pretraining for Domain-Specific Foundation Models

  • 将视觉推理融入预训练,让模型从指令中学习领域特征
  • 在16个遥感和医疗影像任务上达到新SOTA性能
  • 适合需要高精度视觉理解的垂直领域研究者

现代计算机视觉正走向感知、推理与生成相互强化的闭环,但高层推理对底层感知特征的基础学习影响尚未充分探索。本文提出视觉指令预训练(ViTP)新范式,通过在目标下游领域数据上端到端预训练视觉-语言模型,使视觉变压器(ViT)从丰富视觉指令中学习。ViTP引入视觉鲁棒性学习(VRL),促使ViT从稀疏视觉标记中提取稳健且领域相关的特征。在16个具有挑战性的遥感与医学影像基准测试中,ViTP在多样化下游任务上均取得新最优结果。代码已开源:https://github.com/zcablii/ViTP。

原文摘要 · Abstract (English)

Modern computer vision is converging on a closed loop in which perception, reasoning and generation mutually reinforce each other. However, this loop remains incomplete: the top-down influence of high-level reasoning on the foundational learning of low-level perceptual features is not yet underexplored. This paper addresses this gap by proposing a new paradigm for pretraining foundation models in downstream domains. We introduce Visual insTruction Pretraining (ViTP), a novel approach that directly leverages reasoning to enhance perception. ViTP embeds a Vision Transformer (ViT) backbone within a Vision-Language Model and pretrains it end-to-end using a rich corpus of visual instruction data curated from target downstream domains. ViTP is powered by our proposed Visual Robustness Learning (VRL), which compels the ViT to learn robust and domain-relevant features from a sparse set of visual tokens. Extensive experiments on 16 challenging remote sensing and medical imaging benchmarks demonstrate that ViTP establishes new state-of-the-art performance across a diverse range of downstream tasks. The code is available at https://github.com/zcablii/ViTP.

视觉指令领域模型遥感影像医疗影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。