arXiv:2605.07695cs.CV2026-05

无需训练即可精准编辑眼科手术视频,支持文本驱动的器械替换与流程变更。

OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical Videos

论文配图:OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical Videos
图 1 · 摘自论文原文
  • 利用确定性二阶微分方程反演提取原始视频注意力值,实现无训练编辑。
  • 在去噪阶段注入存储的注意力值,保持眼部解剖结构与时间连贯性。
  • 适用于医学教学数据生成,避免繁琐人工录制和昂贵微调。

高保真手术视频生成能显著提升医学培训与AI发展,但实现精确视频编辑仍具挑战,尤其在受严格解剖与时间约束的手术场景中。本文提出OphEdit,一种无需训练的文本引导眼科手术视频编辑框架。该方法通过确定性二阶常微分方程反演管道,从原视频中提取注意力值(V)张量;在去噪阶段,将这些存储的张量选择性注入条件无分类器引导(CFG)分支,从而在严格保留眼部精细解剖结构的同时,无缝实现文本驱动的语义修改。临床评估表明,相较于自然域视频编辑器,OphEdit在器械更换、手术流程变化等复杂操作上展现出更优的结构保真度与时间一致性。本工作首次将无训练视频编辑应用于眼科手术领域,提供了一种无需大量手动录制或昂贵模型微调的可扩展方案,用于生成多样且带标注的医学数据集。代码与提示词详见https://github.com/ophedit/OphEdit。

原文摘要 · Abstract (English)

High-fidelity surgical video generation can greatly improve medical training and the development of AI, adapting these generative models for precise video editing remains a formidable challenge. Modifying surgical attributes, such as instrument tissue interactions or procedural phases is challenging due to the strict anatomical and temporal constraints. In this paper, we propose OphEdit, a novel training-free framework for the text-guided editing of ophthalmic surgical videos. Our approach leverages a deterministic second-order ODE inversion pipeline to capture Attention Value (V) tensors from the original video. By selectively injecting these stored tensors into the conditional Classifier-Free Guidance (CFG) branch during the denoising phase, OphEdit rigorously preserves the intricate anatomical geometry of the eye while seamlessly mapping text-driven semantic modifications onto the video stream. Clinical evaluations demonstrates that OphEdit effectively handles complex surgical transformations, such as instrument swaps and procedural variations, with superior structural fidelity and temporal consistency compared to natural-domain video editors. Our work represents the first application of training-free video editing in the ophthalmic surgical domain, offering a scalable solution for generating diverse, annotated medical datasets without the need for exhaustive manual recording or costly model fine-tuning. The code and prompts can be accessed at https://github.com/ophedit/OphEdit

视频编辑眼科手术无训练医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。