为医疗领域视觉语言模型研究制定分类与报告规范
Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare
- 提出VLM研究的分类框架,区分不同应用阶段
- 建立涵盖评估、数据、论文写作的完整报告标准
- 提供检查清单,提升研究可复现性与发表质量
视觉语言模型(VLM)在医疗领域的开发、适配与应用具有高度复杂性,亟需明确且标准化的报告规范。由于相关研究涵盖从全新VLM构建、领域微调到直接使用预训练模型进行诊断预测等多种范式,传统机器学习报告标准难以适用。本文主张重构现有评估与报告指南,以适应多阶段VLM研究,并兼顾开发者易用性与可复现性要求。为此,我们提出一个VLM研究分类框架,并据此制定涵盖性能评估、数据披露和论文撰写建议的报告标准。所有指南按分类组织,最后提供一份整合检查清单,助力社区统一科研质量标准,推动VLM在医疗领域的规范落地。
原文摘要 · Abstract (English)
The intricate and multifaceted nature of vision language model (VLM) development, adaptation, and application necessitates the establishment of clear and standardized reporting protocols, particularly within the high-stakes context of healthcare. Defining these reporting standards is inherently challenging due to the diverse nature of studies involving VLMs, which vary significantly from the development of all new VLMs or finetuning for domain alignment to off-the-shelf use of VLM for targeted diagnosis and prediction tasks. In this position paper, we argue that traditional machine learning reporting standards and evaluation guidelines must be restructured to accommodate multiphase VLM studies; it also has to be organized for intuitive understanding of developers while maintaining rigorous standards for reproducibility. To facilitate community adoption, we propose a categorization framework for VLM studies and outline corresponding reporting standards that comprehensively address performance evaluation, data reporting protocols, and recommendations for manuscript composition. These guidelines are organized according to the proposed categorization scheme. Lastly, we present a checklist that consolidates reporting standards, offering a standardized tool to ensure consistency and quality in the publication of VLM-related research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。