arXiv:2605.23192cs.CV2026-05

通过智能选帧解决视频编辑中的遮挡问题,提升稳定性与一致性。

Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing

论文配图:Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing
图 1 · 摘自论文原文
  • 从结构完整性、物理稳定性、语义清晰度三方面自动选关键帧。
  • 在多个基准测试中实现更精准、无闪烁的视频编辑效果。
  • 适合需要鲁棒性视频编辑的场景,如动态物体处理或复杂遮挡环境。

视频编辑近年来借助基于扩散模型的生成方法取得显著进展,可实现从自然语言指令出发的多样化物体级操作。然而,现有方法在遮挡、视角变化和快速物体运动下表现不佳,因不可靠的视觉观测导致定位不准、时间闪烁和编辑不一致。本文指出,缺乏可靠视觉锚点是遮挡鲁棒性视频编辑的根本瓶颈。为此,我们提出一种遮挡感知的物理-语义关键帧选择框架,可自动识别下游编辑的最佳锚点帧。该方法从三个互补角度评估候选帧:结构完整性(避免观察截断)、循环一致性跟踪稳定性(衡量物理可靠性)以及基于视觉-语言的属性可见性(确保语义清晰)。所选关键帧通过双向追踪生成密集时空掩码,作为辅助监督信号用于扩散模型视频编辑主干网络。通过将遮挡处理从显式重建转为可靠锚点选择,本框架在无需人工标注的情况下实现精确且时序一致的编辑。大量实验在挑战性视频编辑基准上验证了方法的有效性与高质量表现。

原文摘要 · Abstract (English)

Video editing has recently achieved remarkable progress with diffusion-based generative models, enabling diverse object-level manipulations from natural language instructions. However, existing methods often struggle under occlusion, viewpoint changes, and fast object motion, where unreliable visual observations lead to inaccurate localization, temporal flickering, and inconsistent edits. In this work, we identify the absence of reliable visual anchors as a fundamental bottleneck in occlusion-robust video editing. To address this issue, we propose an occlusion-aware physics-semantic keyframe selection framework that automatically identifies an optimal anchor frame for downstream editing. Specifically, our method evaluates candidate frames from three complementary perspectives: structural completeness for avoiding truncated observations, cycle-consistent tracking stability for measuring physical reliability, and vision-language-based attribute visibility for ensuring semantic clarity. The selected keyframe is then propagated through bidirectional tracking to generate dense spatiotemporal masks, which are used as auxiliary supervision for a diffusion-based video editing backbone. By transforming occlusion handling from explicit reconstruction into reliable anchor selection, our framework enables precise and temporally consistent editing without requiring manual annotations. Extensive experiments on challenging video editing benchmarks demonstrate the effectiveness and high-quality performance of our method.

视频编辑扩散模型关键帧选择遮挡处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。