让骨科手术机器人自主决策,精准执行膝关节截骨。
ArthroCut: Autonomous Policy Learning for Robotic Bone Resection in Knee Arthroplasty
- 用多模态数据训练视觉语言模型,融合术前影像与术中实时感知。
- 在7次实验中平均成功率86%,显著优于现有方法。
- 适合追求高精度、可解释性手术自动化的医疗机器人研究者。
尽管外科手术机器人快速商业化,其自主性和实时决策能力仍受限。为此,我们提出ArthroCut,一种自主策略学习框架,将膝关节置换手术机器人从辅助执行升级为情境感知的动作生成。ArthroCut在自建的21例完整病例数据集(含23,205对RGB-D图像)上微调Qwen-VL主干网络,整合术前CT/MR、术中NDI骨骼与器械追踪、RGB-D手术视频、机器人状态及文本意图等多模态信息。该方法基于两类互补的标记序列:术前影像标记(PIT)用于编码患者解剖结构和规划截骨面,时间对齐手术标记(TAST)用于融合实时视觉、几何与运动学证据,并在语法与安全约束下生成可解释的动作语义。在膝关节假体台架实验中,7次试验平均成功率达86%,显著优于同协议训练的强基线。消融实验证明TAST是可靠性主要驱动因素,而PIT提供关键解剖基准,二者结合实现最稳定的多平面执行。结果表明,将术前几何与时间对齐的术中感知对齐,并转化为受控的标记化动作,是实现稳健、可解释的骨科手术机器人自主性的有效路径。
原文摘要 · Abstract (English)
Despite rapid commercialization of surgical robots, their autonomy and real-time decision-making remain limited in practice. To address this gap, we propose ArthroCut, an autonomous policy learning framework that upgrades knee arthroplasty robots from assistive execution to context-aware action generation. ArthroCut fine-tunes a Qwen--VL backbone on a self-built, time-synchronized multimodal dataset from 21 complete cases (23,205 RGB--D pairs), integrating preoperative CT/MR, intraoperative NDI tracking of bones and end effector, RGB--D surgical video, robot state, and textual intent. The method operates on two complementary token families -- Preoperative Imaging Tokens (PIT) to encode patient-specific anatomy and planned resection planes, and Time-Aligned Surgical Tokens (TAST) to fuse real-time visual, geometric, and kinematic evidence -- and emits an interpretable action grammar under grammar/safety-constrained decoding. In bench-top experiments on a knee prosthesis across seven trials, ArthroCut achieves an average success rate of 86% over the six standard resections, significantly outperforming strong baselines trained under the same protocol. Ablations show that TAST is the principal driver of reliability while PIT provides essential anatomical grounding, and their combination yields the most stable multi-plane execution. These results indicate that aligning preoperative geometry with time-aligned intraoperative perception and translating that alignment into tokenized, constrained actions is an effective path toward robust, interpretable autonomy in orthopedic robotic surgery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。