构建大规模多模态手术数据与模型,提升外科AI泛化能力
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
- 整合6大临床专科、18项手术任务的多模态数据,统一标注标准
- 涵盖超598万条对话,支持理解、推理、规划等复杂任务
- 引入分层推理标注,增强模型在复杂场景下的可解释性
外科智能有望提升手术安全性和一致性,但现有AI框架多为特定任务设计,跨术式和机构泛化能力差。尽管多模态大模型在医疗领域展现强跨任务能力,其在手术领域的进展受限于缺乏大规模、系统化的多模态数据。为此,我们提出Surg$Σ$,一套面向外科智能的大规模多模态数据与基础模型体系。核心是Surg$Σ$-DB,一个整合开源数据、院内临床资料及网络来源数据的多模态数据基础,采用统一模式,提升标签一致性和数据标准化水平。Surg$Σ$-DB覆盖6个临床专科、多种手术类型,包含18项实用手术任务的图像与视频级标注,涵盖理解、推理、规划与生成,规模达598万+对话。除常规多模态对话外,还引入分层推理标注,提供更丰富的语义线索,支持复杂手术场景的深层理解。我们通过基于Surg$Σ$-DB构建的最新外科基础模型验证了其有效性,证明大规模多模态标注、统一语义设计与结构化推理标注对提升跨任务泛化与可解释性的实际价值。
原文摘要 · Abstract (English)
Surgical intelligence has the potential to improve the safety and consistency of surgical care, yet most existing surgical AI frameworks remain task-specific and struggle to generalize across procedures and institutions. Although multimodal foundation models, particularly multimodal large language models, have demonstrated strong cross-task capabilities across various medical domains, their advancement in surgery remains constrained by the lack of large-scale, systematically curated multimodal data. To address this challenge, we introduce Surg$Σ$, a spectrum of large-scale multimodal data and foundation models for surgical intelligence. At the core of this framework lies Surg$Σ$-DB, a large-scale multimodal data foundation designed to support diverse surgical tasks. Surg$Σ$-DB consolidates heterogeneous surgical data sources (including open-source datasets, curated in-house clinical collections and web-source data) into a unified schema, aiming to improve label consistency and data standardization across heterogeneous datasets. Surg$Σ$-DB spans 6 clinical specialties and diverse surgical types, providing rich image- and video-level annotations across 18 practical surgical tasks covering understanding, reasoning, planning, and generation, at an unprecedented scale (over 5.98M conversations). Beyond conventional multimodal conversations, Surg$Σ$-DB incorporates hierarchical reasoning annotations, providing richer semantic cues to support deeper contextual understanding in complex surgical scenarios. We further provide empirical evidence through recently developed surgical foundation models built upon Surg$Σ$-DB, illustrating the practical benefits of large-scale multimodal annotations, unified semantic design, and structured reasoning annotations for improving cross-task generalization and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。