arXiv:2509.06033cs.CV2025-09

用通用视觉语言模型自动解读血检报告,帮患者理解医学数据

Analysis of Blood Report Images Using General Purpose Vision-Language Models

  • 用三款通用视觉语言模型分析100张血检图片,按临床问题提问
  • 模型回答与标准答案的语义相似度达0.78以上,表现良好
  • 适合开发患者端医疗辅助工具,提升健康素养

血检报告的准确解读对健康管理至关重要,但普通人常难以理解,导致焦虑和遗漏问题。本文探索通用视觉语言模型(VLMs)在自动分析血检图像方面的潜力。在包含100张多样化血检图像的数据集上,对比评估了Qwen-VL-Max、Gemini 2.5 Pro和Llama 4 Maverick三款模型。每张报告均针对其内容提出临床相关问题,并通过Sentence-BERT计算模型回答与标准答案的语义相似度进行评估。结果表明,通用VLMs在血检图像理解方面具有实际应用前景,能从图像中直接提供清晰解释,有助于提升健康素养,缓解对复杂医学信息的理解障碍。本研究为未来可靠、可及的AI医疗辅助应用奠定基础。尽管结果积极,仍需注意样本量有限带来的局限性。

原文摘要 · Abstract (English)

The reliable analysis of blood reports is important for health knowledge, but individuals often struggle with interpretation, leading to anxiety and overlooked issues. We explore the potential of general-purpose Vision-Language Models (VLMs) to address this challenge by automatically analyzing blood report images. We conduct a comparative evaluation of three VLMs: Qwen-VL-Max, Gemini 2.5 Pro, and Llama 4 Maverick, determining their performance on a dataset of 100 diverse blood report images. Each model was prompted with clinically relevant questions adapted to each blood report. The answers were then processed using Sentence-BERT to compare and evaluate how closely the models responded. The findings suggest that general-purpose VLMs are a practical and promising technology for developing patient-facing tools for preliminary blood report analysis. Their ability to provide clear interpretations directly from images can improve health literacy and reduce the limitations to understanding complex medical information. This work establishes a foundation for the future development of reliable and accessible AI-assisted healthcare applications. While results are encouraging, they should be interpreted cautiously given the limited dataset size.

医疗影像视觉语言模型健康科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。