首个评估多模态大模型在青少年脊柱侧弯自管中应用的研究
Adapting and Evaluating Multimodal Large Language Models for Adolescent Idiopathic Scoliosis Self-Management: A Divide and Conquer Framework
- 分而治之框架评估视觉问答、知识理解与患者教育三任务
- 检索增强生成显著提升知识任务表现,但脊柱畸形定位准确率仅0.55
- 脊柱关键点提示对不同模型效果不一,当前模型难胜任个性化护理
本研究首次系统评估了多模态大语言模型(MLLMs)在青少年特发性脊柱侧弯(AIS)自我管理中的应用。我们构建了约3,000张前后位X光片及对应诊断文本的数据集,通过‘分而治之’框架,包含视觉问答、领域知识评估和患者教育咨询三个任务,评估五种MLLMs。研究发现,现有模型在解读复杂脊柱影像和理解AIS护理知识方面存在明显局限。为此,我们提出基于脊柱关键点的提示方法,并构建AIS知识库用于检索增强生成(RAG)。结果显示,不同架构对视觉提示的响应差异显著,而RAG大幅提升了知识任务表现。然而,模型在精准识别脊柱畸形位置(最高准确率0.55)和方向(最高准确率0.13)方面仍严重不足,表明当前MLLMs距离实现个性化AIS护理助手仍有巨大差距。
原文摘要 · Abstract (English)
This study presents the first comprehensive evaluation of Multimodal Large Language Models (MLLMs) for Adolescent Idiopathic Scoliosis (AIS) self-management. We constructed a database of approximately 3,000 anteroposterior X-rays with diagnostic texts and evaluated five MLLMs through a `Divide and Conquer' framework consisting of a visual question-answering task, a domain knowledge assessment task, and a patient education counseling assessment task. Our investigation revealed limitations of MLLMs' ability in interpreting complex spinal radiographs and comprehending AIS care knowledge. To address these, we pioneered enhancing MLLMs with spinal keypoint prompting and compiled an AIS knowledge base for retrieval augmented generation (RAG), respectively. Results showed varying effectiveness of visual prompting across different architectures, while RAG substantially improved models' performances on the knowledge assessment task. Our findings indicate current MLLMs are far from capable in realizing personalized assistant in AIS care. The greatest challenge lies in their abilities to obtain accurate detections of spinal deformity locations (best accuracy: 0.55) and directions (best accuracy: 0.13).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。