arXiv:2602.12004cs.AI2026-02

用语言模型评估生成医学图像的临床语义准确性

CSEval: A Framework for Evaluating Clinical Semantics in Text-to-Image Generation

  • 利用语言模型比对图像与文本提示的临床语义一致性
  • 能发现传统方法忽略的解剖位置和病灶错位问题
  • 适合医疗AI安全评估,助力生成模型临床落地

文本到图像生成在医疗领域被广泛用于数据增强和教学。但现有评估方法主要关注图像真实感或多样性,未能衡量生成图像是否准确反映预期的临床语义(如解剖位置和病理特征)。本文提出临床语义评估框架CSEval,利用语言模型评估生成图像与输入提示之间的临床语义对齐程度。实验表明,CSEval能识别其他指标遗漏的语义不一致,并与专家判断高度相关。该框架为现有评估方法提供了可扩展且具有临床意义的补充,支持生成模型在医疗场景中的安全应用。

原文摘要 · Abstract (English)

Text-to-image generation has been increasingly applied in medical domains for various purposes such as data augmentation and education. Evaluating the quality and clinical reliability of these generated images is essential. However, existing methods mainly assess image realism or diversity, while failing to capture whether the generated images reflect the intended clinical semantics, such as anatomical location and pathology. In this study, we propose the Clinical Semantics Evaluator (CSEval), a framework that leverages language models to assess clinical semantic alignment between the generated images and their conditioning prompts. Our experiments show that CSEval identifies semantic inconsistencies overlooked by other metrics and correlates with expert judgment. CSEval provides a scalable and clinically meaningful complement to existing evaluation methods, supporting the safe adoption of generative models in healthcare.

医学图像生成评估临床语义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。