用记忆化动作分块让机器人学会精准扫描组织表面。
Memorized action chunking with Transformers: Imitation learning for vision-based tissue surface scanning
- 用历史图像预测未来动作序列,解决长程依赖问题。
- 50次示范即达60%-80%成功率,优于基线模型。
- 适合需要精细运动的手术辅助场景,如癌症切除。
光学感知技术正用于癌症手术中,确保癌变组织完全清除。虽然点状评估有潜力,但自动化大范围扫描可实现整体组织采样。然而,这类任务因长时程依赖和精细运动要求而困难。为此,我们提出记忆化动作分块与Transformer结合的方法(MACT),利用过去图像序列作为历史信息,预测近未来动作序列。同时采用混合时空位置嵌入以促进学习。在多种仿真环境中,MACT在轮廓扫描和区域扫描上均显著优于基线模型。真实测试中,仅需50条示范轨迹,便在所有扫描任务中达到60%-80%的成功率。结果表明,MACT是手术场景下自适应扫描的有前景模型。
原文摘要 · Abstract (English)
Optical sensing technologies are emerging technologies used in cancer surgeries to ensure the complete removal of cancerous tissue. While point-wise assessment has many potential applications, incorporating automated large area scanning would enable holistic tissue sampling. However, such scanning tasks are challenging due to their long-horizon dependency and the requirement for fine-grained motion. To address these issues, we introduce Memorized Action Chunking with Transformers (MACT), an intuitive yet efficient imitation learning method for tissue surface scanning tasks. It utilizes a sequence of past images as historical information to predict near-future action sequences. In addition, hybrid temporal-spatial positional embeddings were employed to facilitate learning. In various simulation settings, MACT demonstrated significant improvements in contour scanning and area scanning over the baseline model. In real-world testing, with only 50 demonstration trajectories, MACT surpassed the baseline model by achieving a 60-80% success rate on all scanning tasks. Our findings suggest that MACT is a promising model for adaptive scanning in surgical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。