首个同步脑电与第一视角视频的大型数据集,助力行为理解研究
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
- 构建首个多模态同步数据集,融合第一视角视觉与32通道脑电
- 跨被试跨环境验证下动作识别准确率达66.70%
- 适合脑机接口、行为分析与多模态学习研究者使用
脑机接口(BCI),特别是脑电图(EEG)与人工智能(AI)的结合,在从神经信号中解码人类认知与行为方面展现出巨大潜力。随着多模态AI模型的发展,前所未有的可能性涌现。本文提出EgoBrain——全球首个大规模、时间对齐的多模态数据集,同步记录40名参与者在29类日常活动中长达61小时的32通道EEG与第一视角视频,建立以人为中心的行为分析新范式。我们构建了多模态学习框架,融合脑电与视觉信息进行动作理解,在跨被试与跨环境挑战下取得66.70%的动作识别准确率。EgoBrain为统一的多模态与第一视角脑机接口铺平道路,连接神经信号与第一人称感知。数据集与代码已公开:https://huggingface.co/datasets/ut-vision/EgoBrain 与 https://github.com/ut-vision/EgoBrain。
原文摘要 · Abstract (English)
The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has shown tremendous promise in decoding human cognition and behavior from neural signals. In particular, the rise of multimodal AI models have brought new possibilities that have never been imagined before. Here, we present EgoBrain -- the world's first large-scale, temporally aligned multimodal dataset that synchronizes first-person (egocentric) vision and EEG of human brain over extended periods of time, establishing a new paradigm for human-centered behavior analysis. This dataset comprises 61 hours of synchronized 32-channel EEG recordings and first-person video from 40 participants engaged in 29 categories of daily activities. We then developed a multimodal learning framework to fuse EEG and vision for action understanding, validated across both cross-subject and cross-environment challenges, achieving an action recognition accuracy of 66.70%. EgoBrain paves the way toward a unified framework for multimodal and egocentric brain-computer interfaces, bridging neural signals and first-person perception. Our dataset and code are publicly available at: https://huggingface.co/datasets/ut-vision/EgoBrain and https://github.com/ut-vision/EgoBrain .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。