无需标注数据,自动识别手术视频中的关键物体并保持时间连贯性。
Slot-BERT: Self-supervised Object Discovery in Surgical Video
- 采用双向长程注意力机制,在隐空间中学习对象中心表示。
- 在多个真实手术数据集上超越现有无监督方法,表现更优。
- 可零样本迁移至不同手术类型和数据库,适合临床部署。
以对象为中心的槽注意力是一种强大的无监督学习框架,可用于生成结构化且可解释的表示,支持对对象与动作的推理,尤其适用于手术视频。尽管传统视频对象中心方法依赖递归处理以提高效率,但难以维持长视频所需的长期时间连贯性。而完全并行处理整个视频虽增强时间一致性,却带来显著计算开销,不适用于医疗设备硬件。我们提出 Slot-BERT,一种双向长程模型,在保持鲁棒时间连贯性的同时,于隐空间中学习对象中心表示。该方法可无缝扩展至任意长度的长视频。新颖的槽对比损失进一步通过增强槽正交性减少冗余,提升表示解耦。我们在腹腔、胆囊切除术及胸腔手术的真实世界视频数据集上评估了 Slot-BERT。结果表明,其在无监督训练下超越当前最先进对象中心方法,在多种领域表现优异。此外,还展示了对不同外科专业与数据库数据的高效零样本域适应能力。
原文摘要 · Abstract (English)
Object-centric slot attention is a powerful framework for unsupervised learning of structured and explainable representations that can support reasoning about objects and actions, including in surgical videos. While conventional object-centric methods for videos leverage recurrent processing to achieve efficiency, they often struggle with maintaining long-range temporal coherence required for long videos in surgical applications. On the other hand, fully parallel processing of entire videos enhances temporal consistency but introduces significant computational overhead, making it impractical for implementation on hardware in medical facilities. We present Slot-BERT, a bidirectional long-range model that learns object-centric representations in a latent space while ensuring robust temporal coherence. Slot-BERT scales object discovery seamlessly to long videos of unconstrained lengths. A novel slot contrastive loss further reduces redundancy and improves the representation disentanglement by enhancing slot orthogonality. We evaluate Slot-BERT on real-world surgical video datasets from abdominal, cholecystectomy, and thoracic procedures. Our method surpasses state-of-the-art object-centric approaches under unsupervised training achieving superior performance across diverse domains. We also demonstrate efficient zero-shot domain adaptation to data from diverse surgical specialties and databases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。