arXiv:2511.10060cs.CVcs.AI2025-11AAAI被引 1

用高斯表示学习提升医疗动作评估精度与效率

Multivariate Gaussian Representation Learning for Medical Action Evaluation

  • 将动作建模为自适应3维高斯令牌,捕捉时空动态特征
  • 在CPREval-6k数据集上达92.1%准确率,仅需10%计算量
  • 适合需要高精度动作分析的医疗视觉研究者

医学视觉中的细粒度动作评估面临数据集不完整、精度要求严苛及快速动作时空建模不足等挑战。为此,我们构建了包含6,372个专家标注视频、22个临床标签的多视角、多标签基准数据集CPREval-6k。基于此,提出GaussMedAct,一种基于多元高斯编码的框架,通过自适应时空表示学习推进医疗运动分析。多元高斯表示将联合运动投影至时序缩放的多维空间,并分解为自适应3维高斯令牌,作为语义保留的运动单元。该方法通过各向异性协方差建模保持运动语义,同时对时空噪声具有鲁棒性。混合空间编码采用笛卡尔与向量双流策略,有效利用关节与骨骼特征。在基准测试中实现92.1%的Top-1准确率,推理实时,较基线提升5.9%准确率,仅消耗10%浮点运算量。跨数据集实验验证了方法在鲁棒性上的优势。

原文摘要 · Abstract (English)

Fine-grained action evaluation in medical vision faces unique challenges due to the unavailability of comprehensive datasets, stringent precision requirements, and insufficient spatiotemporal dynamic modeling of very rapid actions. To support development and evaluation, we introduce CPREval-6k, a multi-view, multi-label medical action benchmark containing 6,372 expert-annotated videos with 22 clinical labels. Using this dataset, we present GaussMedAct, a multivariate Gaussian encoding framework, to advance medical motion analysis through adaptive spatiotemporal representation learning. Multivariate Gaussian Representation projects the joint motions to a temporally scaled multi-dimensional space, and decomposes actions into adaptive 3D Gaussians that serve as tokens. These tokens preserve motion semantics through anisotropic covariance modeling while maintaining robustness to spatiotemporal noise. Hybrid Spatial Encoding, employing a Cartesian and Vector dual-stream strategy, effectively utilizes skeletal information in the form of joint and bone features. The proposed method achieves 92.1% Top-1 accuracy with real-time inference on the benchmark, outperforming baseline by +5.9% accuracy with only 10% FLOPs. Cross-dataset experiments confirm the superiority of our method in robustness.

医疗视觉动作评估高斯表示时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。