arXiv:2603.19920cs.CV2026-03

解决手术室多视角分割不一致问题,实现无需标定的高精度全景分割。

PanORama: Multiview Consistent Panoptic Segmentation in Operating Rooms

  • 在主干网络中直接建模跨视角特征交互,一次前向传播实现多视角一致性。
  • 在MM-OR和4D-OR数据集上达到70%以上全景质量(PQ),超越现有方法。
  • 无需相机参数标定,可泛化至任意未见视角组合,适合真实手术场景。

手术室环境杂乱、动态且高度遮挡,可靠的时空理解对复杂手术流程中的情境感知至关重要。从稀疏多视角图像中实现全景分割面临根本挑战,因部分视角可见性受限常导致跨相机误判。为此,我们提出PanORama,首个从设计上保证多视角一致性的手术室全景分割方法。通过在单次前向传播中于主干网络内建模跨视角特征交互,视图一致性自然涌现,而非依赖后处理优化。我们在MM-OR和4D-OR数据集上进行评估,实现超过70%的全景质量(PQ)表现,并超越此前最先进方法。重要的是,PanORama无需相机标定参数,可在推理时泛化至任意未见相机视角配置。该方法显著提升多视角分割性能,进而增强手术室中的空间理解能力,为外科感知与辅助开辟新可能。代码将在论文接受后公开。

原文摘要 · Abstract (English)

Operating rooms (ORs) are cluttered, dynamic, highly occluded environments, where reliable spatial understanding is essential for situational awareness during complex surgical workflows. Achieving spatial understanding for panoptic segmentation from sparse multiview images poses a fundamental challenge, as limited visibility in a subset of views often leads to mispredictions across cameras. To this end, we introduce PanORama, the first panoptic segmentation for the operating room that is multiview-consistent by design. By modeling cross-view interactions at the feature level inside the backbone in a single forward pass, view consistency emerges directly rather than through post-hoc refinement. We evaluate on the MM-OR and 4D-OR datasets, achieving >70% Panoptic Quality (PQ) performance, and outperforming the previous state of the art. Importantly, PanORama is calibration-free, requiring no camera parameters, and generalizes to unseen camera viewpoints within any multiview configuration at inference time. By substantially enhancing multiview segmentation and, consequently, spatial understanding in the OR, we believe our approach opens new opportunities for surgical perception and assistance. Code will be released upon acceptance.

全景分割手术室感知多视角一致性无标定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。