EgoPlay根据事件触发指令,精准编辑第一人称视频的前后片段。
EgoPlay: Event-Triggered Video Editing for Egocentric Streams

- 端到端联合学习事件识别与视频编辑,支持多事件和负向提示。
- 在Ego4D上编辑质量、视觉质量和背景一致性提升16%以上。
- 仅用不到一半显存,实现实时流式推理,适合穿戴设备使用。
我们提出EgoPlay,一种基于事件触发的单目第一人称视频编辑器,通过在Ego4D数据集构建的事件条件数据上微调预训练的V2V扩散变换器实现。给定一段单目视频和形如“当X发生时,执行Y”的事件触发提示,EgoPlay可推断事件发生时机,保留事件前帧,并仅对事件后延续部分应用编辑。不同于将事件检测与编辑分离的流水线,EgoPlay在单一端到端模型中联合学习事件识别、时间约束与像素级编辑,同时支持负向和多事件提示。为此,我们构建了包含10.6万对剪辑-提示的大型数据集,涵盖正向触发、伪造负向触发及多事件提示。采用事件触发监督训练双向视频扩散编辑器,并推导出适用于分块流式推理的因果变体。我们还引入事件感知评估协议,分别衡量事件后编辑质量、事件前保留效果与误触发鲁棒性。在Ego4D基准上,EgoPlay显著优于当前最优的指令驱动第一人称视频编辑基线EgoEdit,各项指标相对提升达17.7%、16.9%、16.4%;相比基于视觉语言模型引导的检测-编辑基线,分别提升15.7%、14.5%、13.5%,且显存消耗不足其一半。
原文摘要 · Abstract (English)
We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion transformer on event-conditioned data built primarily from Ego4D. Given a monocular video and an event-triggered prompt of the form "when X happens, do Y," EgoPlay infers whether and when event X occurs, preserves pre-event frames, and applies edit Y only to the post-event continuation. Rather than cascading a separate event detector with an editor, EgoPlay learns event recognition, temporal restraint, and pixel-level editing jointly in a single end-to-end model, while also handling negative and multi-event prompts. To support this, we construct a large-scale dataset of 106K event-triggered clip-prompt pairs spanning positive triggers, fabricated-trigger negatives, and multi-event prompts. We then train a bidirectional video diffusion editor with event-triggered supervision and derive a causal variant for chunk-by-chunk streamable inference. We further introduce an event-aware evaluation protocol that separately measures post-trigger editing quality, pre-trigger preservation, and false-trigger robustness. On the Ego4D benchmark, EgoPlay substantially outperforms EgoEdit, the state-of-the-art instruction-based egocentric video editing baseline, with relative gains of 17.7%, 16.9%, and 16.4% in editing quality, visual quality, and background consistency. It also surpasses a VLM-guided detector-editor baseline by 15.7%, 14.5%, and 13.5% on the same metrics, while using less than half the GPU memory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。