arXiv:2607.01973cs.CVcs.AI2026-07

测试16个视觉语言模型在医学图像质量评估中的可靠性,发现像素化和文本信息会显著影响判断。

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

论文配图:Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias
图 1 · 摘自论文原文
  • 用7种失真类型、5级严重度在MediMeta-C数据集上零样本测试16个VLMs
  • 像素化导致评分平均下降20.58%,最高达34.4%;机构声誉提升评分17.15%
  • 模型对隐私保护与上下文偏见敏感,适合关注医疗AI可靠性的研究者

视觉语言模型(VLMs)在病理描述、报告生成和视觉问答等医疗任务中应用日益广泛。医学图像质量评估(MIQA)通过判断图像是否符合临床决策标准,支持诊断准确性和患者安全。利用VLM自动化MIQA可减轻工作负担,但在图像退化或文本上下文影响判断的真实场景下,其行为仍需深入探究。本研究在MediMeta-C数据集上,零样本评估16个VLMs在七种失真类型和五级严重度下的表现。结果表明:像素化导致评分平均下降20.58%(OCT最高达-34.4%),亮度影响微弱(-0.81%);嵌入空间位移与评分变化相关。同家族模型间评分相关性为0.67–0.83,部分模型在受损乳腺钼靶上评分上升最高达+31%。文本属性显著影响评分:机构声望提升+17.15%,设备年龄降低-14.7%。最大变化为InternVL-8B上升95.62%、MedGemma下降37.7%。当前VLMs在医学图像质量评估中存在局限。像素化作为隐私保护手段却降低性能,揭示隐私与可靠性间的权衡。对上下文元数据的敏感性说明模型缺乏客观性,也暴露元数据是隐私与偏见来源。隐私保护与客观评估是医疗应用中相互关联的要求。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answering. Medical Image Quality Assessment (MIQA) supports diagnostic accuracy and patient safety by determining whether images meet the standards required for clinical decision-making. Automating MIQA with VLMs may reduce workload, but their behavior under real-world conditions, where images may be degraded or textual context may affect judgments, should be further explored before deployment. We benchmark VLMs on medical image quality using the MediMeta-C dataset zero-shot across seven corruption types and five severity levels. We evaluate sensitivity to degradation patterns, the effect of corruptions on embedding geometry, and whether textual attributes (demographics, expertise, infrastructure, institution) alter scores. Across 16 VLMs and seven modalities, pixelation produced the largest score reductions (mean -20.58%, up to -34.4% for OCT), whereas brightness had limited effect (-0.81%). Embedding displacement was associated with score changes. Same-family models showed correlations of 0.67-0.83; some produced increases up to +31% for corrupted mammography. Textual attributes affected scores: institutional prestige raised them +17.15%, and equipment age lowered them -14.7%. The largest changes were +95.62% (InternVL-8B) and -37.7% (MedGemma). Current VLMs show limitations for medical image quality assessment. Pixelation, a privacy-preserving transformation, reduces performance, indicating a trade-off between patient privacy and reliability. Sensitivity to contextual metadata indicates limited objectivity and marks metadata as a privacy and bias source. Privacy protection and objective quality assessment are related requirements for use.

医学图像视觉语言模型可靠性评估隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。