通过时空关联建模,提升手术视频三元组识别的一致性与准确性。
TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition

- 融合空间、关系与时间线索,统一建模三元组内部依赖和跨时序关系。
- 在CholecT45和ProstaTD上分别提升AP_IVT 5.1%和7.8%,TCER降低超36%。
- 适合需要高时序一致性推理的医疗视觉理解任务,如手术分析与辅助决策。
理解复杂手术场景需识别多个相互依赖的实体(如器械、动作、目标),并保持其跨时间的关系一致性。现有三元组识别方法难以统一建模帧内标签依赖与帧间时序语义。为此,我们提出一个统一框架,整合空间、关系与时间线索,实现鲁棒的手术三元组识别。首先通过多尺度编码器提取类别特定的空间先验;随后由多尺度类激活图引导的关系提取模块(MS-CAMRE)优化这些先验,捕捉静态共现模式与动态上下文依赖。此外,双向时序-关系融合注意力(BTRFA)模块协调时序与关系表示,实现连贯的时序推理。我们还引入新评估指标三元组一致性误差率(TCER),量化模型维持因果与语义一致性的能力。在CholecT45和ProstaTD数据集上的大量实验表明,本方法达到最先进性能,分别提升AP_IVT 5.1%和7.8%;根据TCER,相对减少超过36%和25%,验证了框架在时序-关系联合推理中的有效性。
原文摘要 · Abstract (English)
Understanding complex surgical scenes requires recognizing multiple interdependent entities, such as instruments, actions, and targets, while maintaining their relational consistency across time. Existing surgical triplet recognition methods struggle to jointly model intra-frame label dependencies and inter-frame temporal semantics in a unified manner. To address these limitations, we propose a unified framework that integrates spatial, relational, and temporal cues for robust surgical triplet recognition. Specifically, class-specific spatial priors are first extracted through a multi-scale encoder. These priors are then refined by a Label Correlation Modeling module with multi-scale class activation map-guided relational extraction (MS-CAMRE), enabling the model to capture both static co-occurrence patterns and dynamic contextual dependencies among triplet components. Furthermore, a Bidirectional Temporal-Relational Fusion Attention (BTRFA) module harmonizes temporal and relational representations to achieve coherent temporal reasoning. We also introduce a new evaluation metric, the Triplet Consistency Error Rate (TCER), which quantitatively measures the model's ability to preserve causal and semantic consistency across triplets. Extensive experiments on the CholecT45 and ProstaTD datasets show that our method achieves state-of-the-art performance, improving AP_IVT by 5.1 percent and 7.8 percent, respectively. Moreover, according to TCER, our approach achieves relative reductions of more than 36 percent and 25 percent on the two datasets, respectively, demonstrating the effectiveness of our framework in temporal-relational co-reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。