通过视觉语言闭环增强,实现医学影像精确定位与诊断。
Reinforced Correlation Between Vision and Language for Precise Medical AI Assistant
- 构建视觉与语言双向强化机制,动态优化特征表示。
- 在2000万图像-掩码-描述数据上训练,细胞分割提升23.5%。
- 适用于9种模态165项临床任务,泛化性强且可外推新病种。
医学AI助手在疾病诊断、医学图像分析和报告生成中面临多模态内容精度不足及真实场景验证缺失的挑战。本文提出全栈式医疗AI助手RCMed,通过层次化视觉-语言对齐,在输入与输出端均提升多模态一致性,实现精准解剖定位、病变定界与可靠诊断。其自增强相关机制使视觉特征驱动语言上下文,语言语义引导像素级注意力,形成闭环优化。引入颜色区域描述策略,将解剖结构转化为语义丰富的文本,学习跨尺度的形状-位置-文本关系。在2000万图像-掩码-描述三元组上训练,RCMed在不规则病灶上下文建模与微弱解剖边界识别方面达到当前最优表现,在9种模态下的165项临床任务中性能领先,显现出23.5%相对提升。其强视觉-语言对齐能力支持卓越泛化,在20种临床重要癌种的外部验证中表现优异,涵盖多项新任务。该工作展示了集成多模态模型捕捉细微模式的能力,推动复杂场景下接近人类水平的解读,助力以人为本的AI医疗发展。
原文摘要 · Abstract (English)
Medical AI assistants support doctors in disease diagnosis, medical image analysis, and report generation. However, they still face significant challenges in clinical use, including limited accuracy with multimodal content and insufficient validation in real-world settings. We propose RCMed, a full-stack AI assistant that improves multimodal alignment in both input and output, enabling precise anatomical delineation, accurate localization, and reliable diagnosis through hierarchical vision-language grounding. A self-reinforcing correlation mechanism allows visual features to inform language context, while language semantics guide pixel-wise attention, forming a closed loop that refines both modalities. This correlation is enhanced by a color region description strategy, translating anatomical structures into semantically rich text to learn shape-location-text relationships across scales. Trained on 20 million image-mask-description triplets, RCMed achieves state-of-the-art precision in contextualizing irregular lesions and subtle anatomical boundaries, excelling in 165 clinical tasks across 9 modalities. It achieved a 23.5% relative improvement in cell segmentation from microscopy images over prior methods. RCMed's strong vision-language alignment enables exceptional generalization, with state-of-the-art performance in external validation across 20 clinically significant cancer types, including novel tasks. This work demonstrates how integrated multimodal models capture fine-grained patterns, enabling human-level interpretation in complex scenarios and advancing human-centric AI healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。