用统一视觉表征提升医疗视频中服药依从性识别准确率
AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding
- 构建基于LLaVA的统一视觉隐空间,实现医患视频与语言的预对齐
- 在TB服药视频数据集上,性能比现有模型高3.1%~3.54%
- 适合医疗AI、智能健康管理及可解释性研究者参考
慢性病如糖尿病、高血压、结核病等需严格服药以控制病情。依从性常受患者行为、照护支持、医疗成本等因素影响。本文提出AdCare-VLM,一种基于LLaVA的多模态大模型,通过引入统一视觉隐空间并预先对齐,提升基于患者视频的服药依从性问答能力。使用806段由临床专家标注的结核病服药监控视频进行微调,构建了包含正例、负例和模糊案例的医学依从性VQA数据集LLM-TB-VQA。模型识别出面部清晰度、药物可见性、饮水动作及吞咽行为等视觉特征与医学概念的关联,增强视觉-语言对齐与交互。实验显示,该方法在预训练、常规及低秩适配(LoRA)配置下,相比LLaVA-V1.5与Chat-UniVi等参数高效微调模型,绝对提升3.1%至3.54%。消融实验与注意力图可视化验证了方法有效性,提升了可解释性。
原文摘要 · Abstract (English)
Chronic diseases, including diabetes, hypertension, asthma, HIV-AIDS, epilepsy, and tuberculosis, necessitate rigorous adherence to medication to avert disease progression, manage symptoms, and decrease mortality rates. Adherence is frequently undermined by factors including patient behavior, caregiver support, elevated medical costs, and insufficient healthcare infrastructure. We propose AdCare-VLM, a specialized LLaVA-based multimodal large vision language model (LVLM) by introducing a unified visual latent space with pre-alignment to facilitate visual question answering (VQA) concerning medication adherence through patient videos. We employ a private dataset comprising 806 custom-annotated tuberculosis (TB) medication monitoring videos, which have been labeled by clinical experts, to fine-tune the model for adherence pattern detection. We present LLM-TB-VQA, a detailed medical adherence VQA dataset that encompasses positive, negative, and ambiguous adherence cases. Our method identifies correlations between visual features, such as the clear visibility of the patient's face, medication, water intake, and the act of ingestion, and their associated medical concepts in captions. This facilitates the integration of aligned visual-linguistic representations and improves multimodal interactions. Experimental results indicate that our method surpasses parameter-efficient fine-tuning (PEFT) enabled VLM models, such as LLaVA-V1.5 and Chat-UniVi, with absolute improvements ranging from 3.1% to 3.54% across pre-trained, regular, and low-rank adaptation (LoRA) configurations. Comprehensive ablation studies and attention map visualizations substantiate our approach, enhancing interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。