arXiv:2603.00878cs.CV2026-03被引 2

用多成员时间注意力精准捕捉康复训练中的微小动作,提升评估精度。

MMTA: Multi Membership Temporal Attention for Fine-Grained Stroke Rehabilitation Assessment

  • 每帧可同时关注多个时间窗口,提升动作边界敏感度。
  • 在中风康复数据集上,视频与可穿戴设备输入分别提升1.3和1.6的编辑得分。
  • 单阶段统一架构,适合临床与居家场景,计算开销低。

为支持康复过程中反复评估,自动化评估日常活动中的能力需对治疗视频进行高精度细粒度动作分割。现有时间动作分割模型难以捕捉亚秒级微动作,且易模糊快速阶段转换,限制运动恢复的可靠评估。本文提出多成员时间注意力(MMTA),一种用于细粒度康复评估的高分辨率时序变换器。不同于标准时序注意力每层仅赋予每帧单一注意力上下文,MMTA 允许每帧在同一层内关注多个局部归一化的时间注意力窗口。通过特征空间重叠解析融合这些并行时序视图,在过渡区域保留竞争性局部上下文,同时通过逐层传播实现长程推理。该方法在不增加深度或多阶段优化的前提下提升了边界敏感度。MMTA 在统一单阶段架构中支持视频与可穿戴惯性传感器(IMU)输入,适用于临床与家庭环境。实验显示,其在 StrokeRehab 数据集上相比全局注意力变换器,视频输入提升 1.3、IMU 输入提升 1.6 的编辑得分,且在 50Salads 上进一步提升 3.3。消融实验证明性能提升源于多成员时间视图,而非架构复杂度,为资源受限的康复评估提供实用方案。

原文摘要 · Abstract (English)

To empower the iterative assessments involved during a person's rehabilitation, automated assessment of a person's abilities during daily activities requires temporally precise segmentation of fine-grained actions in therapy videos. Existing temporal action segmentation (TAS) models struggle to capture sub-second micro-movements while retaining exercise context, blurring rapid phase transitions and limiting reliable downstream assessment of motor recovery. We introduce Multi-Membership Temporal Attention (MMTA), a high-resolution temporal transformer for fine-grained rehabilitation assessment. Unlike standard temporal attention, which assigns each frame a single attention context per layer, MMTA lets each frame attend to multiple locally normalized temporal attention windows within the same layer. We fuse these concurrent temporal views via feature-space overlap resolution, preserving competing local contexts near transitions while enabling longer-range reasoning through layer-wise propagation. This increases boundary sensitivity without additional depth or multi-stage refinement. MMTA supports both video and wearable IMU inputs within a unified single-stage architecture, making it applicable to both clinical and home settings. MMTA consistently improves over the Global Attention transformer, boosting Edit Score by +1.3 (Video) and +1.6 (IMU) on StrokeRehab while further improving 50Salads by +3.3. Ablations confirm that performance gains stem from multi-membership temporal views rather than architectural complexity, offering a practical solution for resource-constrained rehabilitation assessment.

动作分割康复评估时序建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。