arXiv:2606.13332cs.CV2026-06

提出首个聚焦手术室细粒度多角色动作的评估基准,推动手术场景理解发展。

OR-Action: Multi-Role Video Understanding with Fine-Grained Actions

论文配图:OR-Action: Multi-Role Video Understanding with Fine-Grained Actions
图 1 · 摘自论文原文
  • 构建基于双视角视频的细粒度动作分类体系,通过场景图状态变化蒸馏生成密集动作标签。
  • 现有图神经网络方法在时序建模上表现不佳,新提出的纯视觉时序模型显著超越基线。
  • 创新多视图特征对齐策略,提升单视角下多角色动作识别性能,减少对穿戴式摄像头依赖。

手术室活动的细粒度理解可实现流程感知辅助,但受限于环境杂乱、遮挡和传感不足。当前主流方法采用场景图作为可解释的交互表示,然而将帧级关系预测转化为时序连续的细粒度动作仍具挑战,缺乏显式时序建模。为此,我们首次在公开可用的主客观双视角手术室数据集上构建了以动作为中心的评估基准,定义细粒度多角色动作分类体系,并通过从真实场景图状态变化中蒸馏生成密集动作段。实验表明,即使引入图神经网络进行显式时序建模,现有场景图预测方法仍难以捕捉时序结构。因此,我们提出一种仅依赖视觉信息的时序模型,在使用全部主视角视频输入时显著优于图基方法。进一步提出一种新颖的多视图到单视图特征对齐策略,有效提升单视角下的多角色动作识别性能,降低了对大量主视角视频采集的需求。基准与代码将在论文接受后发布。

原文摘要 · Abstract (English)

Fine-grained understanding of operating room (OR) activity could enable workflow-aware assistance, yet remains difficult due to clutter, occlusions, and limited sensing. The prevailing approach to model this environment is scene graphs as an interpretable representation of OR interactions. Converting their frame-wise relational predictions into temporally extended, fine-grained actions however, is challenging without explicit temporal modeling. To enable a principled temporal evaluation of current OR understanding methods, we introduce the first action-centric benchmark built on a publicly available ego-exocentric OR dataset by defining a fine-grained, multi-role action taxonomy and generating dense action segments via distillation from ground-truth scene graph state changes. Experiments on this benchmark show that current scene graph prediction methods struggle to model temporal structure, even when adding explicit modeling through Graph Neural Networks. We therefore introduce a vision-only temporal model that outperforms graph-based methods significantly when using all available egocentric video as input. Building on this model we also introduce a novel multi- to single-view feature alignment strategy that improves single-view performance on multi-role action recognition, mitigating the need for extensive egocentric video capture. Benchmark and code will be released upon acceptance.

视频理解手术室分析多角色动作时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。