arXiv:2608.20069cs.CV2026-08

用少数据少计算训练出能生成兽医X光报告的高效专家模型

V-REX: Efficient Specialist VLM Training for Veterinary X-Rays

论文配图:V-REX: Efficient Specialist VLM Training for Veterinary X-Rays
图 1 · 摘自论文原文
  • 重构VLM全流程,改进文本分词与定位机制
  • 参数量仅为通用模型几分之一,诊断报告生成效果更优
  • 适合医疗领域小样本专家模型开发,无需依赖大模型

尽管通用视觉语言模型(VLM)训练成本高昂,普遍认为打造领域专家需微调越来越大的基础模型。我们发现,在兽医放射学中这一假设并不成立。通过重新设计整个VLM流程——从文本分词、预训练到定位与推理——我们证明,通过精心工程化可构建出性能超越更大基础模型的模型,且无需额外数据。本方法引入了新的生成式预训练与定位策略,显著提升训练效率与数据利用率。仅使用当代通用模型极小比例的参数、数据与算力,我们开发出首个能够为兽医放射影像生成诊断报告的VLM,该模型在该任务上表现远超开源基础模型。

原文摘要 · Abstract (English)

While generalist VLMs are expensive to train, creating domain experts is widely assumed to require fine-tuning increasingly large foundation models. We show that, in veterinary radiology, this assumption is misguided. By rethinking the entire VLM pipeline - from text tokenisation and pre-training to grounding and inference - we demonstrate that careful engineering can yield models that outperform much larger foundation models from scratch, without relying on any other data. Our approach introduces new strategies for generative pre-training and grounding that improve training efficiency, increasing data utilisation and downstream performance. Using only a fraction of the parameters, data, and compute of contemporary generalist models, we develop the first VLM capable of generating diagnostic reports for veterinary radiographs, surpassing open foundation models on this task by significant margin.

兽医影像视觉语言模型高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。