arXiv:2412.17595cs.CVcs.AI2024-12被引 3

融合振动信号提升单目胶囊内镜深度与运动估计精度

V$^2$-SfMLearner: Learning Monocular Depth and Ego-motion for Multimodal Wireless Capsule Endoscopy

  • 引入振动信号与视觉信息联合建模,抑制体内振动干扰
  • 在自监督框架下实现更鲁棒的深度与运动估计,性能优于纯视觉方法
  • 适用于临床胶囊机器人,支持实时诊断,适合医学影像领域应用

深度学习可从胶囊内镜视频中预测深度图和胶囊自身运动,辅助三维场景重建与病灶定位。然而,胃肠道内胶囊碰撞引起的振动扰动会污染训练数据。现有方法仅依赖视觉信息,忽略了振动等辅助信号对降噪的潜力。为此,我们提出V²-SfMLearner,一种融合振动信号的多模态方法,用于单目胶囊内镜的深度与运动估计。构建了包含振动与视觉信号的多模态胶囊内镜数据集,所提AI方案采用无监督方法,通过视觉-振动信号融合有效消除振动扰动。具体设计了振动网络分支与傅里叶融合模块,实现振动噪声检测与抑制。该融合框架兼容主流纯视觉算法。在多模态数据集上的大量验证表明,其性能与鲁棒性显著优于纯视觉方法。无需大型外部设备,该方法具备集成至临床胶囊机器人的潜力,提供实时可靠的消化道检查工具,研究结果展示出临床应用前景,有望提升医生诊断能力。

原文摘要 · Abstract (English)

Deep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization. However, the collisions of the capsule endoscopies within the gastrointestinal tract cause vibration perturbations in the training data. Existing solutions focus solely on vision-based processing, neglecting other auxiliary signals like vibrations that could reduce noise and improve performance. Therefore, we propose V$^2$-SfMLearner, a multimodal approach integrating vibration signals into vision-based depth and capsule motion estimation for monocular capsule endoscopy. We construct a multimodal capsule endoscopy dataset containing vibration and visual signals, and our artificial intelligence solution develops an unsupervised method using vision-vibration signals, effectively eliminating vibration perturbations through multimodal learning. Specifically, we carefully design a vibration network branch and a Fourier fusion module, to detect and mitigate vibration noises. The fusion framework is compatible with popular vision-only algorithms. Extensive validation on the multimodal dataset demonstrates superior performance and robustness against vision-only algorithms. Without the need for large external equipment, our V$^2$-SfMLearner has the potential for integration into clinical capsule robots, providing real-time and dependable digestive examination tools. The findings show promise for practical implementation in clinical settings, enhancing the diagnostic capabilities of doctors.

胶囊内镜多模态学习深度估计医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。