无需粗略分割图,通过跨帧信息提升视频抠图的物体识别与边缘精度。
Object-Aware Video Matting with Cross-Frame Guidance
- 利用跨帧信息聚合物体级特征,增强像素级细节表达
- 仅需初始粗略掩码,在多个数据集上达到顶尖性能
- 适合需要精准动态抠图的视频编辑与合成场景
近年来,无trimap方法在人体视频抠图中受到越来越多关注,因其表现优异。然而,这些方法仍缺乏确定性的前景-背景提示,导致难以在时间上一致地识别和定位前景目标,并提取细微细节。本文提出一种无需trimap的物体感知视频抠图框架OAVM,能够感知不同物体,实现前景物体联合识别与边缘细节优化。具体地,我们设计了跨帧引导的物体引导修正与精炼模块(OGCR),将物体级实例信息融入像素级细节特征,促进二者协同。此外,提出序列前景融合增强策略,丰富序列场景,提升网络对物体的区分能力。在多个主流合成与真实世界基准上的大量实验表明,本方法仅需初始粗略掩码即可达到当前最优性能。代码与模型将公开。
原文摘要 · Abstract (English)
Recently, trimap-free methods have drawn increasing attention in human video matting due to their promising performance. Nevertheless, these methods still suffer from the lack of deterministic foreground-background cues, which impairs their ability to consistently identify and locate foreground targets over time and mine fine-grained details. In this paper, we present a trimap-free Object-Aware Video Matting (OAVM) framework, which can perceive different objects, enabling joint recognition of foreground objects and refinement of edge details. Specifically, we propose an Object-Guided Correction and Refinement (OGCR) module, which employs cross-frame guidance to aggregate object-level instance information into pixel-level detail features, thereby promoting their synergy. Furthermore, we design a Sequential Foreground Merging augmentation strategy to diversify sequential scenarios and enhance capacity of the network for object discrimination. Extensive experiments on recent widely used synthetic and real-world benchmarks demonstrate the state-of-the-art performance of our OAVM with only an initial coarse mask. The code and model will be available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。