arXiv:2605.12650cs.CV2026-05被引 1

用临床对齐评分优化医学图像生成,减少幻觉并提升真实感。

CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis

论文配图:CRAFT: Clinical Reward-Aligned Finetuning for Medical Image Synthesis
图 1 · 摘自论文原文
  • 基于临床评分设计奖励机制,融合多模态模型增强生成合理性。
  • 在四类医学影像上显著提升临床对齐度,低对齐尾部降低超20%。
  • 适合医疗图像生成、临床可信度评估及医学AI模型改进研究者。

基础扩散模型能生成逼真的自然图像,但应用于医学影像仍具挑战。由于标注数据有限,现有方法易产生类似幻觉且临床上不合理的合成图像;而传统评估指标如FID或Inception Score无法衡量图像与病理相关标准的对齐程度。本文提出临床对齐评分(CAS),一种基于基础模型的代理指标,从视觉保真度以外的四个互补维度评估生成图像。基于CAS,我们提出临床奖励对齐微调(CRAFT)框架,通过标签条件提示增强、临床检查清单和可微奖励优化,将多模态大语言模型与视觉-语言模型中的医学知识迁移至生成过程。在四种不同模态下,CRAFT在CAS得分和下游分类性能上均优于强基线。除平均CAS提升外,其将低于真实图像参考阈值的低对齐尾部降低了5.5%-34.7个百分点,跨数据集平均相对减少20.4%。结果表明生成幻觉显著减少,且经外部评估者验证、结构化检查清单审计、记忆分析及针对CheXpert的盲测医生偏好实验支持。

原文摘要 · Abstract (English)

Foundation diffusion models can generate photorealistic natural images, but adapting them to medical imaging remains challenging. In medical adaptation, limited labeled data can exacerbate hallucination-like and clinically implausible synthesis, while existing metrics such as FID or Inception Score do not quantify per-image alignment with pathology-relevant criteria. We introduce the Clinical Alignment Score (CAS), a foundation-model-based proxy for clinical alignment that evaluates generated images along four complementary dimensions beyond visual fidelity. Building on CAS, we propose Clinical Reward-Aligned Finetuning (CRAFT), a reward-based adaptation framework that transfers medical knowledge from multimodal large language models and vision-language models through label-conditioned prompt enrichment, clinical checklists, and differentiable reward optimization. Across four diverse modalities, CRAFT improves CAS and downstream classification performance over strong adaptation baselines. Beyond average CAS gains, CRAFT reduces the empirical low-alignment tail below a real-image reference threshold by 5.5-34.7% points relative to the strongest baseline, corresponding to a 20.4% average relative reduction across datasets. These results indicate fewer hallucination-like generations under CAS, and are corroborated by out-of-family evaluator evaluation, structured checklist auditing, memorization analysis, and a blinded physician preference study on CheXpert.

医学图像扩散模型奖励机制临床对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。