arXiv:2603.23757cs.CV2026-03

通过关注身体关节动态,提升癫痫视频检测的跨人泛化能力

Learning Cross-Joint Attention for Generalizable Video-Based Seizure Detection

  • 以关节为中心提取视频片段,抑制背景干扰
  • 学习跨关节注意力,捕捉发作时肢体协同运动模式
  • 在未见受试者上表现优于主流方法,适合真实场景部署

从长期临床视频中实现自动化癫痫检测可大幅减少人工审阅时间并支持实时监测。然而,现有基于视频的方法常因背景偏差和依赖个体外观特征而难以泛化到未见受试者。本文提出一种以关节为中心的注意力模型,仅聚焦身体动态以增强跨受试者泛化性。对每个视频片段,检测身体关节并提取关节中心片段,有效抑制背景上下文。这些片段采用视频视觉变换器(ViViT)进行标记化,并学习跨关节注意力,建模肢体间时空交互,捕捉癫痫发作特有的协调运动模式。大量跨受试者实验表明,该方法在未见受试者上持续优于最先进的基于CNN、图神经网络和Transformer的方法。

原文摘要 · Abstract (English)

Automated seizure detection from long-term clinical videos can substantially reduce manual review time and enable real-time monitoring. However, existing video-based methods often struggle to generalize to unseen subjects due to background bias and reliance on subject-specific appearance cues. We propose a joint-centric attention model that focuses exclusively on body dynamics to improve cross-subject generalization. For each video segment, body joints are detected and joint-centered clips are extracted, suppressing background context. These joint-centered clips are tokenized using a Video Vision Transformer (ViViT), and cross-joint attention is learned to model spatial and temporal interactions between body parts, capturing coordinated movement patterns characteristic of seizure semiology. Extensive cross-subject experiments show that the proposed method consistently outperforms state-of-the-art CNN-, graph-, and transformer-based approaches on unseen subjects.

癫痫检测视频分析注意力机制跨人泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。