arXiv:2605.11265cs.CVcs.AI2026-05中稿 · 29th International…

通过纹理注意力自监督适配,提升手术图像密集预测的泛化能力。

DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction

论文配图:DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction
图 1 · 摘自论文原文
  • 基于纹理中心的槽注意力机制学习不变视觉结构
  • 无需标注即可适应目标域,显著提升跨分布性能
  • 适合需要强泛化能力的微创与机器人手术视觉系统

外科计算机视觉中的密集预测任务(如分割和手术区域预测)可为腹腔镜与机器人手术提供重要指导。然而,由于训练数据难以覆盖部署时的多样性,模型常面临分布偏移问题,导致泛化性能差。本文提出DenseTRF,一种基于纹理中心注意力的自监督表示适配框架。该方法利用槽注意力学习捕捉不变视觉结构的纹理感知表示,并在无监督条件下将这些表示适配至目标分布,显著提升对领域偏移的鲁棒性。框架通过条件化密集预测与模型融合策略实现。在多个外科手术场景下的实验表明,DenseTRF在跨分布泛化能力上优于当前最先进的分割模型及测试分布适配方法。

原文摘要 · Abstract (English)

Dense prediction tasks in surgical computer vision, such as segmentation and surgical zone prediction, can provide valuable guidance for laparoscopic and robotic surgery. However, these models often suffer from distribution shifts, as training datasets rarely cover the variability encountered during deployment, leading to poor generalization. We propose DenseTRF, a self-supervised representation adaptation framework based on texture-centric attention. Our method leverages slot attention to learn texture-aware representations that capture invariant visual structures. By adapting these representations to the target distribution without supervision, DenseTRF significantly improves robustness to domain shifts. The framework is implemented through conditioning dense prediction on slot attention and model merging strategies. Experiments across multiple surgical procedures demonstrate improved cross-distribution generalization in comparison to state-of-the-art segmentation models and test-distribution adaptation methods for dense prediction tasks.

手术视觉自监督密集预测域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。