arXiv:2501.14677cs.CV2025-01CVPR被引 30

无需辅助信息,实现复杂背景下的稳定视频抠图。

MatAnyone: Stable Video Matting with Consistent Memory Propagation

  • 基于自适应记忆融合,跨帧保持主体语义一致。
  • 新数据集+训练策略,显著提升抠图稳定性与细节保真度。
  • 适合真实场景下高精度视频抠图任务,尤其复杂背景

无辅助信息的人体视频抠图方法依赖输入帧本身,常在复杂或模糊背景中表现不佳。为此,我们提出针对目标指定的视频抠图框架 MatAnyone。基于记忆机制,引入区域自适应记忆融合模块,实现前一帧记忆的自适应整合,确保核心区域语义稳定的同时保留对象边界的精细细节。为增强训练鲁棒性,构建了一个更大、更高质量且多样化的视频抠图数据集。此外,提出一种新颖的训练策略,高效利用大规模分割数据,提升抠图稳定性。结合新网络结构、数据集与训练策略,MatAnyone 在多样真实场景中均实现稳健且精确的视频抠图结果,优于现有方法。

原文摘要 · Abstract (English)

Auxiliary-free human video matting methods, which rely solely on input frames, often struggle with complex or ambiguous backgrounds. To address this, we propose MatAnyone, a robust framework tailored for target-assigned video matting. Specifically, building on a memory-based paradigm, we introduce a consistent memory propagation module via region-adaptive memory fusion, which adaptively integrates memory from the previous frame. This ensures semantic stability in core regions while preserving fine-grained details along object boundaries. For robust training, we present a larger, high-quality, and diverse dataset for video matting. Additionally, we incorporate a novel training strategy that efficiently leverages large-scale segmentation data, boosting matting stability. With this new network design, dataset, and training strategy, MatAnyone delivers robust and accurate video matting results in diverse real-world scenarios, outperforming existing methods.

视频抠图记忆机制图像分割鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。