用多模态大模型自动识别癫痫发作的病理动作,效果优于传统模型。
Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology

- 用零样本方式评估大模型对20种癫痫症状的识别能力。
- 在18项特征中13项表现超基准模型,尤其擅长姿势和上下文识别。
- 通过图像裁剪等预处理提升性能,且解释可信度达94.3%。
多模态大语言模型(MLLMs)在识别日常人类行为方面表现出强大能力,但在分析神经疾病中的临床显著不自主运动方面仍缺乏探索。本小规模研究评估了先进MLLMs在自动识别癫痫视频中病理运动方面的潜力。我们对90段临床癫痫发作录像中的20个ILAE定义的半性征特征进行了零样本测试。结果显示,无需任务特定训练,MLLMs在18项特征中的13项上超越了微调的卷积神经网络(CNN)和视觉变换器(ViT)基线模型,尤其在识别显著体位与情境特征方面表现突出,但对细微高频动作识别较弱。针对特定特征的信号增强策略(如面部裁剪、姿态估计、音频降噪)使20项特征中10项性能提升。专家评估显示,94.3%的正确预测案例中,模型生成的解释忠实度达到60%以上,与癫痫专科医生推理一致。这些发现表明,通过针对性预处理,可将通用型MLLMs应用于专业临床视频分析,为可解释、高效的辅助诊断提供新路径。代码已公开于 https://github.com/LinaZhangUCLA/PathMotionMLLM。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have demonstrated robust capabilities in recognizing everyday human activities, yet their potential for analyzing clinically significant involuntary movements in neurological disorders remains largely unexplored. This pilot study evaluates the capability of MLLMs for automated recognition of pathological movements in seizure videos. We assessed the zero-shot performance of state-of-the-art MLLMs on 20 ILAE-defined semiological features across 90 clinical seizure recordings. MLLMs outperformed fine-tuned Convolutional Neural Network (CNN) and Vision Transformer (ViT) baseline models on 13 of 18 features without task-specific training, demonstrating particular strength in recognizing salient postural and contextual features while struggling with subtle, high-frequency movements. Feature-targeted signal enhancement (facial cropping, pose estimation, audio denoising) improved performance on 10 of 20 features. Expert evaluation showed that 94.3 percent of MLLM-generated explanations for correctly predicted cases achieved at least 60 percent faithfulness scores, aligning with epileptologist reasoning. These findings demonstrate the potential of adapting general-purpose MLLMs for specialized clinical video analysis through targeted preprocessing strategies, offering a path toward interpretable, efficient diagnostic assistance. Our code is publicly available at https://github.com/LinaZhangUCLA/PathMotionMLLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。