用可解释的视觉语言模型,让AI像医生一样分析伤口感染并给出理由。
Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning

- 用大模型生成推理链,小模型学习专家级诊断逻辑。
- 在多源伤口数据上达到86.8%准确率,优于主流模型。
- 输出带视觉证据支持的诊断理由,适合临床辅助决策。
从照片评估慢性伤口感染极具挑战,因外观受病因、部位和成像条件影响。以往基于图像的深度学习方法多关注分类,缺乏可解释性,难以支持临床决策。本文提出Infection-Reasoner,一个40亿参数的紧凑型视觉语言推理模型,用于慢性伤口感染分类与推理生成。针对标注推理数据稀缺问题,采用两阶段训练:(1) 推理蒸馏,由GPT-5.1为无标签伤口图像生成思维链,引导小模型Qwen3-VL-4B-Thinking学习特定伤口推理;(2) 在少量标注感染数据上,通过组相对策略优化(Group Relative Policy Optimization)进行强化学习微调。在独立异构伤口数据集上,模型达86.8%准确率、86.4%敏感度、87.1%特异度,优于多个强基线,包括GPT-5.1。推理质量经多模态大模型与伤口专家双重评估:四名MLLM裁判的视觉支持一致率在0.722至0.903之间,专家评审中61.8%的推理为正确,32.4%为部分正确。
原文摘要 · Abstract (English)
Assessing chronic wound infection from photographs is challenging because visual appearance varies across wound etiologies, anatomical locations, and imaging conditions. Prior image-based deep learning methods have mainly focused on classification with limited interpretability, despite the need for evidence-grounded explanations to support point-of-care decision making. We present Infection-Reasoner, a compact 4B-parameter reasoning vision-language model for chronic wound infection classification and rationale generation. To address the scarcity of expert-labeled wound images with reasoning annotations, Infection-Reasoner is trained using a two-stage pipeline: (1) reasoning distillation, in which GPT-5.1 generates chain-of-thought rationales for unlabeled wound images to initialize wound-specific reasoning in a smaller student model (Qwen3-VL-4B-Thinking), and (2) reinforcement learning post-training with Group Relative Policy Optimization on a small labeled infection dataset to refine classification reasoning. On a held-out heterogeneous wound dataset, Infection-Reasoner achieved 86.8\% accuracy, 86.4\% sensitivity, and 87.1\% specificity, outperforming several strong baselines, including GPT-5.1. Rationale quality was further evaluated using both multimodal large language model (MLLM) judges and wound expert review. Across four MLLM judges, visual-support agreement scores ranged from 0.722 to 0.903, while expert review rated 61.8\% of rationales as Correct and 32.4\% as Partially Correct.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。