arXiv:2603.13402cs.CVcs.LG2026-03中稿 · ECCV

让视频生成更懂动作逻辑,用事件信号纠正物体互动错误

Event-Driven Video Generation

  • 引入事件驱动机制,仅在关键交互区域更新潜在表示
  • 在EVD-Bench上提升动作连贯性与支撑关系稳定性,优于基线模型
  • 适合关注视频时序逻辑、物理一致性研究的开发者

当前文本到视频模型虽能生成逼真单帧图像,但在简单交互上仍存在明显错误:物体在接触前移动、动作被跳过、放置物持续漂移或支撑关系断裂。本文认为,传统逐帧去噪方法在每一步都更新所有潜在区域,而提示词仅应激活局部交互。为此提出事件驱动视频生成(EVD),一种轻量级DiT兼容干预方案。通过轻量头预测令牌级事件活跃度,训练损失将其与潜在状态变化对齐;采用带滞后性和早期调度的事件门控采样,在交互形成区域集中应用更新。在EVD-Bench测试中,EVD显著提升人类偏好及VBench动态指标,包括状态保持、空间精度、支撑关系和接触稳定性,同时保持与基线模型相当的外观质量。结果表明,少量事件结构即可修正隐藏在良好帧级外观下的多种交互错误。

原文摘要 · Abstract (English)

Current text-to-video models can make individual frames look convincing while still getting simple interactions wrong: objects move before contact, an intended action is skipped, a placed object keeps drifting, or a support relation breaks. Our starting point is that standard frame-first denoising updates every latent region at every step, even when the prompt implies that only a local interaction should be active. We introduce Event-Driven Video Generation (EVD), a small DiT-compatible intervention that gives the sampler an explicit event signal. A lightweight head predicts token-level event activity; training losses tie that activity to latent state change; and event-gated sampling, with hysteresis and an early-step schedule, applies the update field mainly where an interaction is forming. On EVD-Bench, EVD improves human preference and VBench dynamics for state persistence, spatial accuracy, support relations, and contact stability, while keeping appearance quality comparable to the base model. The results suggest that a modest amount of event structure can correct several interaction failures that otherwise remain hidden behind good frame-level appearance.

视频生成事件驱动时序一致性扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。