基于1090万数据训练的视觉基础模型,实现多模态疼痛自动评估
PainFormer: a Vision Foundation Model for Automatic Pain Assessment
- 用14个任务/数据集联合训练,构建多任务视觉基础模型
- 在BioVid和AI4Pain上超越75种方法,多模态表现最优
- 支持图像、热成像、深度、生理信号等多模态输入,适合临床应用
疼痛是影响大量人群的复杂病症,准确可靠的评估对制定有效管理方案至关重要。自动疼痛评估系统可实现连续监测并辅助决策,减轻患者痛苦、预防功能退化。本文提出PainFormer,一种基于多任务学习的视觉基础模型,在14个任务/数据集上共1090万样本上训练而成。该模型作为嵌入提取器,为Embedding-Mixer(基于Transformer的模块)提供特征表示,完成最终疼痛评估。实验采用行为模态(如RGB、合成热成像、估计深度视频)和生理模态(如ECG、EMG、GSR、fNIRS),验证了PainFormer能从多种输入中提取高质量嵌入。在BioVid和AI4Pain两个数据集上,与文献中75种方法直接对比,无论单模态还是多模态设置均达领先水平,推动通用自动疼痛评估模型的发展。模型架构与权重已开源:https://github.com/GkikasStefanos/PainFormer。
原文摘要 · Abstract (English)
Pain is a manifold condition that impacts a significant percentage of the population. Accurate and reliable pain evaluation for the people suffering is crucial to developing effective and advanced pain management protocols. Automatic pain assessment systems provide continuous monitoring and support decision-making processes, ultimately aiming to alleviate distress and prevent functionality decline. This study introduces PainFormer, a vision foundation model based on multi-task learning principles trained simultaneously on 14 tasks/datasets with a total of 10.9 million samples. Functioning as an embedding extractor for various input modalities, the foundation model provides feature representations to the Embedding-Mixer, a transformer-based module that performs the final pain assessment. Extensive experiments employing behavioral modalities - including RGB, synthetic thermal, and estimated depth videos - and physiological modalities such as ECG, EMG, GSR, and fNIRS revealed that PainFormer effectively extracts high-quality embeddings from diverse input modalities. The proposed framework is evaluated on two pain datasets, BioVid and AI4Pain, and directly compared to 75 different methodologies documented in the literature. Experiments conducted in unimodal and multimodal settings demonstrate state-of-the-art performances across modalities and pave the way toward general-purpose models for automatic pain assessment. The foundation model's architecture (code) and weights are available at: https://github.com/GkikasStefanos/PainFormer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。