用临床结构化奖励提升医疗影像描述准确性,让生成结果更符合真实医学逻辑。
Clinically Structured Surrogate Rewards for Post-SFT Medical Image Captioning

- 设计结构化奖励,匹配图像邻域分布与临床图一致性
- 在六个组合中平均提升事实性5.8%、相关性2.1%、整体表现3.4%
- 适合需要高可信度医疗文本生成的临床应用
医疗影像描述需将异构视觉证据转化为简洁的临床描述,错误的发现、断言状态或解剖关系即使语言流畅也会改变临床意义。序列级策略优化可直接优化完整描述,但常用奖励依赖全局文本相似度、图像-描述匹配或无序概念重叠,忽视了视觉邻域和临床-断言结构。本文提出一种后SFT医疗影像描述的临床结构化代理奖励框架。该框架结合生物医学语义与短程词汇保真度,引入两项结构化奖励:分布式图像邻域对齐(匹配参考与生成描述所诱导的影像库分布)和临床图一致性(对实体、断言状态、类型化关系进行最大权重一对一匹配)。四项奖励在每批内独立归一化,固定权重组合后通过GDPO优化。在ImageCLEFmedical标准与合成标题赛道的组织者评估隐藏测试集上,三种视觉语言骨干网络均实现整体、相关性和事实性超越匹配的SFT基线,平均相对提升分别为3.4%、2.1%和5.8%。消融与成对诊断表明结构化奖励提供互补信号,减少图像邻域差异并提升实体-断言-关系一致性。
原文摘要 · Abstract (English)
Medical image captioning requires translating heterogeneous visual evidence into concise clinical descriptions, where errors in findings, assertion states, or anatomical relations can alter clinical meaning despite surface-level fluency. Sequence-level policy optimization can directly optimize complete captions, but common rewards rely on global text similarity, direct image-caption compatibility, or unordered concept overlap, leaving visual neighborhoods and clinical-claim structure implicit. We propose a clinically structured surrogate reward framework for post-SFT medical image captioning. The framework combines biomedical semantic and short-range lexical fidelity with two structured rewards: distributional image-neighborhood alignment, which matches the medical-image-bank distributions induced by reference and generated captions, and clinical graph consistency, which applies maximum-weight one-to-one matching to entities, assertion states, and typed relations. The four rewards are independently normalized within each rollout group, combined with fixed relative weights, and optimized with GDPO. Across organizer-evaluated hidden test sets for the Standard and Synthetical ImageCLEFmedical Caption tracks and three vision-language backbones, the method improves Overall, Relevance, and Factuality over matched SFT baselines in all six backbone-track combinations, with average relative gains of 3.4%, 2.1%, and 5.8%, respectively. Ablations and paired diagnostics indicate that the structured rewards provide complementary signals, reducing image-neighborhood divergence and improving entity-assertion-relation consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。