SurgVisAgent用多模态大模型实现手术视觉的智能增强
SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
- 基于多模态大模型动态识别图像失真类型和程度
- 支持低光、过曝、模糊、烟雾等多类增强,效果优于单一任务模型
- 适合临床医生和医疗AI研发者参考,提升手术视觉体验
精准的外科手术对患者安全至关重要,已有先进增强算法辅助医生决策。然而,现有算法多针对特定场景的单一任务,难以应对复杂真实环境。为此,我们提出SurgVisAgent,一个基于多模态大语言模型(MLLMs)的端到端智能手术视觉代理。该模型可动态识别内窥镜图像中的失真类别与严重程度,进而执行低光增强、过曝校正、运动模糊消除和烟雾去除等多种任务。为提升手术场景理解能力,我们设计了领域专用先验模型;通过上下文少样本学习与思维链(CoT)推理,实现针对不同失真类型和严重程度的定制化图像增强。此外,我们构建了一个模拟真实手术失真的综合基准,大量实验表明SurgVisAgent超越传统单任务模型,展现出作为统一手术辅助解决方案的巨大潜力。
原文摘要 · Abstract (English)
Precise surgical interventions are vital to patient safety, and advanced enhancement algorithms have been developed to assist surgeons in decision-making. Despite significant progress, these algorithms are typically designed for single tasks in specific scenarios, limiting their effectiveness in complex real-world situations. To address this limitation, we propose SurgVisAgent, an end-to-end intelligent surgical vision agent built on multimodal large language models (MLLMs). SurgVisAgent dynamically identifies distortion categories and severity levels in endoscopic images, enabling it to perform a variety of enhancement tasks such as low-light enhancement, overexposure correction, motion blur elimination, and smoke removal. Specifically, to achieve superior surgical scenario understanding, we design a prior model that provides domain-specific knowledge. Additionally, through in-context few-shot learning and chain-of-thought (CoT) reasoning, SurgVisAgent delivers customized image enhancements tailored to a wide range of distortion types and severity levels, thereby addressing the diverse requirements of surgeons. Furthermore, we construct a comprehensive benchmark simulating real-world surgical distortions, on which extensive experiments demonstrate that SurgVisAgent surpasses traditional single-task models, highlighting its potential as a unified solution for surgical assistance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。