用LoRA实现临床伤口多模态自适应,提升严重并发症检测能力
Cross-Contextual Vision-Language Adaptation with LoRA for Personalized Severe Adverse Event Detection in Clinical Wound Monitoring

- 基于冻结的BiomedCLIP,通过双流LoRA融合临床文本与伤口图像特征
- 在纵向数据上实现91.3%的严重不良事件检测准确率,优于基线模型
- 支持个性化风险识别,适合临床长期伤口监控场景
伤口监测是关键但未充分解决的临床挑战,及时发现感染、组织退化和愈合延迟等严重不良事件(SAEs)可显著改善患者预后。尽管视觉-语言模型(VLMs)具备强大的多模态推理能力,但通常缺乏领域特定的语义基础,难以整合伤口影像与异构临床信息,且对训练分布外案例的检测机制有限。本文提出一种自动伤口监测与SAE检测的多模态框架。该方法利用配对的临床记录与伤口描述,捕捉外观、周围皮肤状况、颜色变化及炎症或愈合进展等视觉特征,通过基于冻结的BiomedCLIP主干网络构建的双流低秩适配(LoRA)框架进行编码。引入跨上下文LoRA融合机制,在不微调整个模型的前提下,实现临床语义与视觉伤口描述间的双向信息交互,生成上下文感知的多模态表示。为识别个性化SAE,提出结合语义匹配、视觉典型性、图文对齐与图象-文本对齐的统一OOD(分布外)检测框架,计算综合评分。为捕捉愈合动态,引入协变量一致性与时间漂移惩罚项,利用随访中伤口特征的变化建模。在临床随访收集的纵向伤口数据集上的实验表明,该框架在伤口愈合评估与SAE检测任务中均表现优异,验证了语义丰富且时序感知的视觉-语言系统在临床伤口监测与早期风险识别中的潜力。
原文摘要 · Abstract (English)
Wound monitoring is a critical yet underserved clinical challenge, where timely identification of severe adverse events (SAEs) such as infection, tissue deterioration, and delayed healing can significantly impact patient outcomes. While vision-language models (VLMs) show strong multimodal reasoning, they often lack domain-specific grounding to integrate wound imagery with heterogeneous clinical information, and provide limited mechanisms for detecting cases that diverge from the training distribution. We present a multimodal framework for automated wound monitoring and SAE detection. Our approach leverages paired clinical notes and wound descriptions capturing visual characteristics such as appearance, surrounding skin condition, color changes, and signs of inflammation or healing progression, encoded through a dual-stream Low-Rank Adaptation (LoRA) framework built on a frozen BiomedCLIP backbone. We introduce a cross-contextual LoRA fusion mechanism enabling information exchange between clinical semantics and visual wound descriptors, producing context-aware multimodal representations without full model fine-tuning. To identify personalized SAEs, we propose a wound-specific out-of-distribution (OOD) detection framework combining semantic matching, visual typicality, caption-text alignment, and caption-visual alignment into a unified SAE (OOD) score. To capture healing dynamics, we incorporate covariate consistency and temporal drift penalties that leverage changes in wound characteristics across visits. Experiments on a longitudinal wound dataset collected through clinical visits show promising performance on both wound healing assessment and SAE detection, highlighting the potential of semantically enriched, temporally aware vision-language systems for clinical wound monitoring and early risk identification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。