构建癫痫发作行为分析数据集与评测体系,推动多模态模型临床应用
Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

- 基于438段视频建立20类癫痫症状密集标注数据集
- 提出七项分层任务评测,发现主流模型在定位与诊断上存在系统性缺陷
- 引入临床可解释报告评估指标,适配医生使用场景
尽管多模态大模型在通用视频理解中表现优异,但对其解析癫痫等不自主、时空演变的病理运动行为的能力仍缺乏验证。为此,我们提出了癫痫症状学套件(S3),一个面向临床的数据库与评测基准,包含438段癫痫发作视频,覆盖20种国际抗癫痫联盟(ILAE)定义的症状特征,共35,000多个密集标注。在此基础上,我们设计了七项分层任务,从低层视觉感知到时序推理、叙事生成及诊断,系统评估多模态大模型。为实现生成报告的临床有效性评估,我们提出癫痫报告质量指数(Seizure-RQI)。11个开源多模态模型的基线测试揭示其在侧别判断、时间定位、症状序列与临床报告准确性上的系统性短板。结果显示,针对癫痫任务微调可显著提升性能,两阶段神经符号框架在癫痫与非癫痫发作分类任务上达到F1=0.96。S3为安全关键医疗视频理解提供了严格评测标准,推动领域自适应多模态智能的发展。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret involuntary, and spatio-temporally evolving pathologic motor behaviors such as seizure semiology remains largely untested. To address this gap, we introduce Seizure-Semiology-Suite, a clinically grounded dataset and benchmark for fine-grained, structured seizure semiology understanding. The dataset includes 438 seizure videos annotated with over 35,000 dense labels covering 20 ILAE-defined semiological features. Building on this dataset, we propose a seven-task hierarchical benchmark that systematically evaluates MLLMs from low-level visual perception to temporal sequencing, narrative report generation, and seizure diagnosis. To enable clinically meaningful evaluation of generated reports, we further introduce the Report Quality Index for Seizure Semiology (Seizure-RQI). Extensive baselines across 11 open-weight MLLMs reveal systematic weaknesses in laterality reasoning, temporal localization, symptom sequencing, and clinically faithful reporting. We show that seizure-specific fine-tuning substantially improves performance across tasks, and that a two-stage neuro-symbolic framework achieves an F1 score of 0.96 on epileptic versus non-epileptic seizure classification. Seizure-Semiology-Suite establishes a rigorous benchmark for evaluating multimodal models in safety-critical medical video understanding and guides the development of clinically reliable, domain-adaptive multimodal intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。