arXiv:2505.07449eess.IVcs.CV2025-05中稿 · MICCAI25被引 7

用自然语言指令生成眼科手术视频,解决数据隐私与标注难题

Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model

  • 基于16万对图文数据,构建端到端生成模型
  • 生成视频在医生评估中达真实可靠水平
  • 适合用于手术流程理解与AI训练场景

眼科手术中,开发能解读手术视频并预测后续操作的AI系统,需大量高质量标注的手术视频,但受隐私和人力成本限制难以获取。文本引导视频生成(T2V)为解决此问题提供新路径,可依据外科医生指令生成手术视频。本文提出Ophora,首个能根据自然语言指令生成眼科手术视频的模型。我们首先设计全面的数据清洗流程,将叙述性眼科手术视频转化为包含超过16万对视频-指令的数据集Ophora-160K;随后提出渐进式视频-指令微调方案,将预训练于自然视频-文本数据集的T2V模型的空间-时间知识迁移至隐私保护的眼科手术视频生成任务。通过定量分析与眼科医生反馈验证,Ophora可生成逼真且可靠的手术视频。我们还验证了其在眼科手术流程理解等下游任务中的潜力。代码已开源。

原文摘要 · Abstract (English)

In ophthalmic surgery, developing an AI system capable of interpreting surgical videos and predicting subsequent operations requires numerous ophthalmic surgical videos with high-quality annotations, which are difficult to collect due to privacy concerns and labor consumption. Text-guided video generation (T2V) emerges as a promising solution to overcome this issue by generating ophthalmic surgical videos based on surgeon instructions. In this paper, we present Ophora, a pioneering model that can generate ophthalmic surgical videos following natural language instructions. To construct Ophora, we first propose a Comprehensive Data Curation pipeline to convert narrative ophthalmic surgical videos into a large-scale, high-quality dataset comprising over 160K video-instruction pairs, Ophora-160K. Then, we propose a Progressive Video-Instruction Tuning scheme to transfer rich spatial-temporal knowledge from a T2V model pre-trained on natural video-text datasets for privacy-preserved ophthalmic surgical video generation based on Ophora-160K. Experiments on video quality evaluation via quantitative analysis and ophthalmologist feedback demonstrate that Ophora can generate realistic and reliable ophthalmic surgical videos based on surgeon instructions. We also validate the capability of Ophora for empowering downstream tasks of ophthalmic surgical workflow understanding. Code is available at https://github.com/uni-medical/Ophora.

视频生成眼科手术文本生成视频数据隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。