MuSS数据集解决视频生成中的连贯叙事与角色一致性难题
MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

- 构建分阶段字幕流水线,先保证单镜头准确再整合全局逻辑
- 引入跨镜头匹配机制,彻底消除主体生成的复制粘贴问题
- 提供电影级叙事评测基准,适合追求长序列生成的研究者
尽管视频基础模型在单镜头生成上表现优异,但真实影视叙事依赖复杂的多镜头编排。当前进展受限于缺乏解决三大核心挑战的数据集:真实叙事逻辑、时空文本-视频对齐冲突,以及主体到视频生成中的‘复制粘贴’缺陷。为此,我们提出MuSS——一个大规模双轨数据集,专为多镜头视频与主体到视频生成设计。数据源自超过3000部电影,明确支持复杂蒙太奇过渡与以主体为中心的叙事。构建过程中,我们首创渐进式字幕流水线,通过先确保单镜头准确性,再强化全局叙事连贯性,消除上下文冲突。关键在于,我们实现跨镜头匹配机制,从根本上杜绝S2V中的复制粘贴捷径。同时,提出电影叙事评测基准,包含视觉-逻辑驱动范式与新型抗复制粘贴方差(ACP-Var)指标,严格评估连续叙事与三维结构一致性。大量实验表明,现有基线在连续叙事逻辑上表现不佳或退化为二维贴纸生成器,而使用MuSS训练的模型达到最先进的叙事有效性与跨镜头身份保留效果。
原文摘要 · Abstract (English)
While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained by the absence of datasets that address three core challenges: authentic narrative logic, spatiotemporal text-video alignment conflicts, and the "copy-paste" dilemma prevalent in Subject-to-Video (S2V) generation. To bridge this gap, we introduce MuSS, a large-scale, dual-track dataset tailored for multi-shot video and S2V generation. Sourced from over 3,000 movies, MuSS explicitly supports both complex montage transitions and subject-centric narratives. To construct this dataset, we pioneer a progressive captioning pipeline that eliminates contextual conflicts by ensuring local shot-level accuracy before enforcing global narrative coherence. Crucially, we implement a cross-shot matching mechanism to fundamentally eradicate the S2V copy-paste shortcut. Alongside the dataset, we propose the Cinematic Narrative Benchmark, featuring a visual-logic-driven paradigm and a novel Anti-Copy-Paste Variance (ACP-Var) metric to rigorously assess continuous storytelling and 3D structural consistency. Extensive experiments demonstrate that while current baselines struggle with continuous narrative logic or degenerate into trivial 2D sticker generators, our MuSS-augmented model achieves state-of-the-art narrative effectiveness and cross-shot identity preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。