对比多种视频大模型在帕金森病远程筛查中的表现,发现不同模型擅长不同任务。
Benchmarking Video Foundation Models for Remote Parkinson's Disease Screening
- 用7种视频大模型评估16项临床任务,分析其在帕金森病筛查中的适用性。
- 最高准确率达80.6%,特异性达90.3%,但敏感性仅43.2%-57.3%。
- 为远程神经监测提供模型与任务匹配的选型指南,适合临床研究者参考。
基于视频的评估为远程帕金森病(PD)筛查提供了可扩展路径。传统方法依赖手工特征模拟临床量表,而近期发展的视频基础模型(VFMs)可在无需特定任务定制的情况下实现表示学习。然而,不同VFMs架构在多样临床任务中的相对有效性尚不明确。本研究基于1,888名参与者(727名帕金森病患者)的新视频数据集,涵盖32,847段视频和16项标准化临床任务,系统评估了七种先进VFMs——包括VideoPrism、V-JEPA、ViViT和VideoMAE——在临床筛查中的鲁棒性。通过冻结嵌入并使用线性分类头评估,结果表明任务相关性高度依赖模型:VideoPrism在捕捉视觉语音动力学(无音频)和面部表情方面表现优异,V-JEPA在上肢运动任务中更优,TimeSformer在指节敲击等节奏任务中仍具竞争力。实验获得的AUC范围为76.4%–85.3%,准确率范围为71.5%–80.6%。尽管高特异性(最高达90.3%)表明其在排除健康个体方面潜力显著,但较低的敏感性(43.2%–57.3%)凸显了需进行任务感知校准及多任务多模态融合的必要性。本工作建立了基于视频大模型的帕金森病筛查基准,并为远程神经监测中模型与任务的选择提供路线图。代码与匿名结构化数据已公开:https://anonymous.4open.science/r/parkinson_video_benchmarking-A2C5
原文摘要 · Abstract (English)
Video-based assessments offer a scalable pathway for remote Parkinson's disease (PD) screening. While traditional approaches rely on handcrafted features mimicking clinical scales, recent advances in video foundation models (VFMs) enable representation learning without task-specific customization. However, the comparative effectiveness of different VFM architectures across diverse clinical tasks remains poorly understood. We present a large-scale systematic study using a novel video dataset from 1,888 participants (727 with PD), comprising 32,847 videos across 16 standardized clinical tasks. We evaluate seven state-of-the-art VFMs -- including VideoPrism, V-JEPA, ViViT, and VideoMAE -- to determine their robustness in clinical screening. By evaluating frozen embeddings with a linear classification head, we demonstrate that task saliency is highly model-dependent: VideoPrism excels in capturing visual speech kinematics (no audio) and facial expressivity, while V-JEPA proves superior for upper-limb motor tasks. Notably, TimeSformer remains highly competitive for rhythmic tasks like finger tapping. Our experiments yield AUCs of 76.4 - 85.3% and accuracies of 71.5 - 80.6%. While high specificity (up to 90.3%) suggests strong potential for ruling out healthy individuals, the lower sensitivity (43.2 - 57.3%) highlights the need for task-aware calibration and integration of multiple tasks and modalities. Overall, this work establishes a rigorous baseline for VFM-based PD screening and provides a roadmap for selecting suitable tasks and architectures in remote neurological monitoring. Code and anonymized structured data are publicly available: https://anonymous.4open.science/r/parkinson\_video\_benchmarking-A2C5
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。