arXiv:2603.21349cs.CV2026-03

用视频变压器识别呼吸窘迫,准确率达81%

Respiratory Status Detection with Video Transformers

  • 用运动后恢复过程构建带标签数据集,以时间顺序标注呼吸状态
  • 提出结合李相对编码与运动引导掩码的视频变换器,F1达0.81
  • 适合临床早期预警系统开发,尤其关注呼吸监测场景

通过视觉观察识别呼吸窘迫是救命的临床技能。医生可发现呼吸恶化早期迹象,为早期干预创造宝贵窗口。本研究评估近期视频变换器能否从视频中识别呼吸窘迫征兆。我们采集了健康志愿者剧烈运动后的恢复视频,利用个体自然恢复过程构建呼吸状态标签数据集。将视频切分为短片段,较早片段对应更严重气促,设计时序排序挑战以评估AI系统识别能力。结果表明,采用李相对编码(LieRE)与运动引导掩码的ViViT编码器,结合基于嵌入的比较策略,可在该任务上取得0.81的F1分数。研究提示现代视频变换器能捕捉呼吸力学的细微变化。

原文摘要 · Abstract (English)

Recognition of respiratory distress through visual inspection is a life saving clinical skill. Clinicians can detect early signs of respiratory deterioration, creating a valuable window for earlier intervention. In this study, we evaluate whether recent advances in video transformers can enable Artificial Intelligence systems to recognize the signs of respiratory distress from video. We collected videos of healthy volunteers recovering after strenuous exercise and used the natural recovery of each participants respiratory status to create a labeled dataset for respiratory distress. Splitting the video into short clips, with earlier clips corresponding to more shortness of breath, we designed a temporal ordering challenge to assess whether an AI system can detect respiratory distress. We found a ViViT encoder augmented with Lie Relative Encodings (LieRE) and Motion Guided Masking, combined with an embedding based comparison strategy, can achieve an F1 score of 0.81 on this task. Our findings suggest that modern video transformers can recognize subtle changes in respiratory mechanics.

视频生成呼吸检测变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。