构建首个消息驱动视频编辑评估数据集,揭示叙事目标如何决定剪辑选择。
MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing

- 基于多消息-多剪辑配对设计,展示同一视频因叙事目标不同而剪辑差异显著。
- 模型在严格标准下远逊于人类,尤其在时间对齐精度上差距明显。
- 引入消息模糊性与上下文丰富度评分,量化叙事难度影响模型表现。
视频编辑本质上由叙事意图驱动:相同素材因表达目标不同而选取不同镜头。现有视频摘要基准将编辑意图简化为单一的、无关消息的显著性概念,无法反映这种多样性。为此,我们提出 extbf{MEDit-Bench},一个数据集与评估基准,将长视频与多个编辑消息及每条消息对应多个专业剪辑配对,证明相同源视频在不同消息下产生显著不同的剪辑结果。我们定义基于时间对齐度量的自动评估协议,并发现以大模型为裁判的偏好判断存在严重位置偏差,不可靠。此外,我们为每条消息标注模糊性与上下文丰富度得分,发现两者均与模型性能负相关,确立消息难度为有意义的分层因子。实验表明,尽管先进多模态大模型在宽松阈值下接近人类时间对齐水平,但在严格标准下仍显著落后;人类感知研究进一步确认巨大质量差距,专业人工剪辑始终优于模型输出。
原文摘要 · Abstract (English)
Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes to convey. Benchmarks for a closely related task, video summarization, reduce editorial intent to a single, message-agnostic notion of saliency and thus do not account for this diversity. For evaluating message-driven video editing, we present \textbf{MEDit-Bench}, a dataset and benchmark, which pairs long-form videos with multiple editing messages and multiple professionally produced edits per message, demonstrating that different messages yield substantially different edits from the same source. We define an automatic evaluation protocol based on temporal alignment metrics, and find that an LLM-as-a-judge preference, a natural proxy for narrative quality, is unreliable for this task due to severe position bias. We additionally annotate each message with ambiguity and contextfulness scores, and show that both dimensions negatively correlate with model performance, establishing message difficulty as a meaningful stratification factor. Experiments with state-of-the-art MLLMs and reinforcement fine-tuned baselines show that while strong models approach human temporal alignment at lenient thresholds, all models fall behind humans at stricter criteria. A human perceptual study further confirms a large quality gap, with professional human edits remaining consistently preferred over model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。