arXiv:2507.07242cs.CV2025-07

用AI自动生成视频实例掩码,提升影视特效制作效率

Automated Video Segmentation Machine Learning Pipeline

  • 通过文本提示实现灵活物体检测
  • 支持帧间一致的分割与跟踪
  • 适合影视特效团队快速部署使用

视觉特效(VFX)制作常受限于耗时且资源密集的遮罩生成。本文提出一种自动化视频分割流水线,可生成时间一致的实例掩码。该方法利用机器学习实现:(1)基于文本提示的灵活物体检测,(2)精细化的逐帧图像分割,(3)鲁棒的视频跟踪以确保时间稳定性。通过容器化部署并采用结构化输出格式,该流水线被艺术家迅速采纳。显著减少人工工作量,加快初步合成稿的生成速度,并提供完整的分割数据,从而提升整体VFX生产效率。

原文摘要 · Abstract (English)

Visual effects (VFX) production often struggles with slow, resource-intensive mask generation. This paper presents an automated video segmentation pipeline that creates temporally consistent instance masks. It employs machine learning for: (1) flexible object detection via text prompts, (2) refined per-frame image segmentation and (3) robust video tracking to ensure temporal stability. Deployed using containerization and leveraging a structured output format, the pipeline was quickly adopted by our artists. It significantly reduces manual effort, speeds up the creation of preliminary composites, and provides comprehensive segmentation data, thereby enhancing overall VFX production efficiency.

视频分割自动化VFX

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。