用扩散变换器让胶囊机器人听懂指令,精准操控胃内任务
CapsDT: Diffusion-Transformer for Capsule Robot Manipulation
- 融合视觉、文本与控制信号的扩散变换器架构
- 在仿真中实现26.25%真实操作成功率,覆盖四类内镜任务
- 专为胃部胶囊机器人设计,适合医疗自动化研究者
视觉-语言-动作(VLA)模型在多个领域展现出巨大潜力,但在内窥镜胶囊机器人——即在消化道内执行操作的微型机器人——中的应用仍属空白。将VLA模型引入内窥镜机器人可实现人机更直观高效的交互,提升诊断准确率与治疗效果。本文提出CapsDT,一种用于胃内胶囊机器人操作的扩散变换器模型,通过处理交错的视觉输入与文本指令,推断出相应的机器人控制信号以完成内镜任务。此外,我们构建了基于机械臂持磁控的胶囊内镜机器人系统,在胃模拟器中完成四类不同难度的任务,并建立了对应的胶囊机器人数据集。在多种机器人任务上的综合评估表明,CapsDT可作为强大的视觉-语言通用模型,在各类内镜任务中均达到当前最优表现,并在真实世界仿真操作中实现26.25%的成功率。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as a prominent research area, showcasing significant potential across a variety of applications. However, their performance in endoscopy robotics, particularly endoscopy capsule robots that perform actions within the digestive system, remains unexplored. The integration of VLA models into endoscopy robots allows more intuitive and efficient interactions between human operators and medical devices, improving both diagnostic accuracy and treatment outcomes. In this work, we design CapsDT, a Diffusion Transformer model for capsule robot manipulation in the stomach. By processing interleaved visual inputs, and textual instructions, CapsDT can infer corresponding robotic control signals to facilitate endoscopy tasks. In addition, we developed a capsule endoscopy robot system, a capsule robot controlled by a robotic arm-held magnet, addressing different levels of four endoscopy tasks and creating corresponding capsule robot datasets within the stomach simulator. Comprehensive evaluations on various robotic tasks indicate that CapsDT can serve as a robust vision-language generalist, achieving state-of-the-art performance in various levels of endoscopy tasks while achieving a 26.25% success rate in real-world simulation manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。