让机器人听懂医生指令,自动规划内镜视角,精准查看鼻腔结构。
EndoNav: Semantic-to-Geometric Grounding for Language-Guided Robotic Endoscopic Examination

- 基于患者特异性解剖结构,将语言指令转化为可视化目标。
- 在3个鼻腔模型上实现87%以上视野重合度,接近医生水平。
- 适合智能手术导航、机器人辅助内镜系统研发人员参考。
在受限解剖空间中进行的微创手术依赖持续的内窥镜可视化。当前机器人内窥镜系统可稳定或调整内镜位置,但缺乏相关上下文以提供有效可视化辅助。本文提出EndoNav,一种基于解剖结构的自然语言框架,可将高层医生指令转化为患者特异性鼻窦解剖结构中的自主内窥镜可视化行为。医生口头指令经转录后,由基于患者特异性解剖场景表示的内窥镜视角代理进行解读。该代理不直接生成机器人运动,而是生成结构化可视化目标,转换为靶点视角与检查轨迹,并通过几何约束的内窥镜运动规划和关节空间控制执行。我们在三个基于CT重建的解剖模型上进行了三阶段鼻窦检查评估。对一具尸体标本,将自主可视化结果与两名住院医师的检查进行对比:EndoNav的平均视野交并比分别为87.04%和84.37%,接近两名医生间的交并比87.44%;同时恢复了92.91%和93.20%的医生观察到的解剖表面。结果表明,将高层解剖指令接地为患者特异性几何目标,并转化为解剖约束下的机器人可视化行为是可行的。
原文摘要 · Abstract (English)
Minimally invasive procedures performed within confined anatomical spaces depend on continuous endoscopic visualization. Current robotic endoscope systems can stabilize or reposition an endoscope, but they do not possess relevant context to provide effective visualization assistance. We present EndoNav, an anatomy-grounded natural-language framework that translates high-level surgeon commands into autonomous endoscopic visualization behaviors within patient-specific sinonasal anatomy. Spoken surgeon commands are transcribed and interpreted by an endoscopic viewpoint agent conditioned on a patient-specific anatomical scene representation. Rather than generating robot motion directly, the viewpoint agent generates structured visualization objectives that are converted into target viewpoints and inspection trajectories, which are then executed through geometry-constrained endoscope motion planning and joint-space control. We evaluate EndoNav using a structured three-pass sinus examination across three CT-derived anatomical models. For one cadaveric specimen, autonomous visualization is compared with sinus examinations performed by two resident surgeons. EndoNav achieved mean visualization IoUs of 87.04% and 84.37% relative to the two surgeon examinations, compared with an inter-surgeon IoU of 87.44%, while recovering 92.91% and 93.20% of surgeon-observed anatomical surfaces, respectively. These results demonstrate the feasibility of grounding high-level anatomical commands into patient-specific geometric objectives and translating them into anatomically constrained robotic visualization behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。