通过双阶段推理提升医学影像报告的视觉准确性,减少幻觉。
Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
- 用临床语言和指标奖励引导生成语义完整报告
- 无需标注或检索,仅靠图像-文本相似度驱动输出
- 适合需要高可信医学报告生成的研究与临床场景
放射科报告生成(RRG)是自动化医疗流程、提升患者评估准确性和减轻医务人员负担的关键步骤。尽管大型医学视觉语言模型(Med-VLMs)取得进展,但生成既视觉对齐又临床准确的报告仍面临挑战。现有方法依赖大规模标注数据、昂贵的任务特定偏好数据或基于检索的知识库,难以有效缓解视觉与语言表征间对齐不良导致的幻觉问题。为此,我们提出VALOR:面向精准放射科报告生成的医学视觉语言模型对齐方法,通过两个互补推理阶段解决视觉幻觉:(1) 临床引导的文本推理,利用可验证的自然语言和临床指标奖励,生成术语精确、语义完整的报告;(2) 自监督视觉推理,借助冻结的领域专家计算输入胸片与生成候选间的图像-文本相似度,转换为归一化优势值,显式引导策略生成视觉对齐输出,无需偏好对、检索数据库或额外标注。在多个基准上的实验表明,VALOR显著提升生成质量与临床准确性,性能超越当前最优医学报告生成模型。
原文摘要 · Abstract (English)
Radiology Report Generation (RRG) is a critical step toward automating healthcare workflows, facilitating accurate patient assessments, and reducing the workload of medical professionals. Despite recent progress in Large Medical Vision-Language Models (Med-VLMs), generating radiology reports that are both visually grounded and clinically accurate remains a significant challenge. Existing approaches often rely on large labeled corpora for pre-training, costly task-specific preference data, or retrieval-based knowledge. However, these strategies do not adequately mitigate hallucinations arising from poor cross-modal alignment between visual and linguistic representations. To address these limitations, we propose VALOR: Visual Alignment of Medical Vision-Language Models for GrOunded Radiology Report Generation, which tackles visual hallucinations through two complementary reasoning stages: (1) Clinically Informed Textual Reasoning guides the model with verifiable natural language and clinical metric rewards to produce semantically complete reports with precise medical terminology. (2) Self-Supervised Visual Reasoning leverages a frozen domain expert to compute image-text similarity scores between the input chest X-ray and generated candidates, converting these into rank-normalized advantages that explicitly steer the policy toward visually grounded outputs, requiring no preference pairs, retrieval databases, or additional annotations. Extensive experiments on multiple benchmarks demonstrate that VALOR substantially improves generation quality, as well as clinical accuracy which are visually grounded, achieving significant performance gains over state-of-the-art medical report generation benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。