用4D空间时间表示让手术AI无需训练就能理解视频中的工具与组织位置变化。
A 4D Representation for Training-Free Agentic Reasoning from Monocular Laparoscopic Video
- 构建显式4D模型,融合点追踪、深度与分割信息,统一时空语义。
- 在134个临床问题上,推理准确率显著提升,实现自然语言与4D空间的精准对齐。
- 仅用现成2D大模型和3D视觉工具,零微调即达成智能推理,适合医疗AI研究者。
时空推理是人工智能在软组织手术中的一项基础能力,为智能辅助系统和自主机器人铺平道路。尽管2D视觉-语言模型在理解手术视频方面展现出越来越大的潜力,但手术场景的空间复杂性表明,显式的4D表示可能有助于推理系统。本文提出一种框架,通过显式4D表示赋予手术代理时空推理能力,使AI系统能够将自然语言推理同时锚定在时间和3D空间中。利用点追踪、深度估计和分割模型,我们构建了一个具有时空一致性工具与组织语义的连贯4D模型。随后,多模态大语言模型(MLLM)作为代理,基于该4D表示生成的工具轨迹等数据进行推理,且无需任何微调。我们在一个包含134个临床相关问题的新数据集上评估了该方法,结果表明,通用推理主干与我们的4D表示相结合,显著提升了时空理解能力,并实现了4D定位。我们证明,无需额外训练,即可从2D MLLM和3D计算机视觉模型“组装”出时空智能。代码、数据与示例见 https://tum-ai.github.io/surg4d/
原文摘要 · Abstract (English)
Spatiotemporal reasoning is a fundamental capability for artificial intelligence (AI) in soft tissue surgery, paving the way for intelligent assistive systems and autonomous robotics. While 2D vision-language models show increasing promise at understanding surgical video, the spatial complexity of surgical scenes suggests that reasoning systems may benefit from explicit 4D representations. Here, we propose a framework for equipping surgical agents with spatiotemporal tools based on an explicit 4D representation, enabling AI systems to ground their natural language reasoning in both time and 3D space. Leveraging models for point tracking, depth, and segmentation, we develop a coherent 4D model with spatiotemporally consistent tool and tissue semantics. A Multimodal Large Language Model (MLLM) then acts as an agent on tools derived from the explicit 4D representation (e.g., trajectories) without any fine-tuning. We evaluate our method on a new dataset of 134 clinically relevant questions and find that the combination of a general purpose reasoning backbone and our 4D representation significantly improves spatiotemporal understanding and allows for 4D grounding. We demonstrate that spatiotemporal intelligence can be "assembled" from 2D MLLMs and 3D computer vision models without additional training. Code, data, and examples are available at https://tum-ai.github.io/surg4d/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。