用语言控制机器人X光机,让医生说句话就能自动定位和成像。
Intelligent Control of Robotic X-ray Devices using a Language-promptable Digital Twin
- 用语言提示的AI模型动态更新患者数字孪生,实现语义理解与视觉反馈。
- 在尸体实验中,84%的指令能端到端成功完成定位与照射。
- 可从任意角度精准定位35个常见结构,误差小于51.68毫米,适合临床实用。
自然语言为控制机器人C臂X光系统提供了一种便捷灵活的接口,使高级功能更易访问。然而,实现语言控制需要专门的AI模型对X光图像进行语义解析以支持推理,而这些模型固定输出限制了语言控制的功能。通过使用可语言提示的通用模型进行X光图像分割,本系统基于稀疏重建的解剖结构持续更新患者数字孪生,支持自主可视化、个性化视角定位及自动准直功能,例如执行‘聚焦下腰椎’等指令。在尸体研究中,用户通过语音命令实现了全身范围内的结构可视化、定位与照射,端到端成功率达84%。事后分析显示,该数字孪生模型可在随机姿态下将35个常用结构定位至51.68毫米以内,实现任意角度下的准确隔离。结果表明,智能机器人X光系统可直接融合医生的表达意图。尽管现有术中X光分析的通用模型存在失效情况,但随着其改进,将推动高度灵活、智能的机器人C臂发展。
原文摘要 · Abstract (English)
Natural language offers a convenient, flexible interface for controlling robotic C-arm X-ray systems, making advanced functionality and controls accessible. However, enabling language interfaces requires specialized AI models that interpret X-ray images to create a semantic representation for reasoning. The fixed outputs of such AI models limit the functionality of language controls. Incorporating flexible, language-aligned AI models prompted through language enables more versatile interfaces for diverse tasks and procedures. Using a language-aligned foundation model for X-ray image segmentation, our system continually updates a patient digital twin based on sparse reconstructions of desired anatomical structures. This supports autonomous capabilities such as visualization, patient-specific viewfinding, and automatic collimation from novel viewpoints, enabling commands 'Focus in on the lower lumbar vertebrae.' In a cadaver study, users visualized, localized, and collimated structures across the torso using verbal commands, achieving 84% end-to-end success. Post hoc analysis of randomly oriented images showed our patient digital twin could localize 35 commonly requested structures to within 51.68 mm, enabling localization and isolation from arbitrary orientations. Our results demonstrate how intelligent robotic X-ray systems can incorporate physicians' expressed intent directly. While existing foundation models for intra-operative X-ray analysis exhibit failure modes, as they improve, they can facilitate highly flexible, intelligent robotic C-arms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。