将视频描述拆解为可独立编辑的多个信息流,提升生成与理解质量。
Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding

- 把视频分成参考、镜头、事件和全局四类信息流,分离处理
- 在多个数据集上降低25%错误率,推理任务性能提升67%
- 显著改善角色一致性、音画对齐和时间控制,适合视频生成研究者
多模态大语言模型正推动视频字幕从描述性输出转向理解与生成的语义接口。然而,主流方法仍将视频视为整体叙述段落,混合视觉、音频与身份信息,导致表示失真且难以扩展,局部修改需全局重写。为此,我们提出多流场景剧本(MTSS)新范式,用因子化与显式关联的场景描述替代整体文本。核心包含两大原则:流分解,将视频解耦为参考、镜头、事件与全局四类互补流;关系锚定,通过显式身份与时间链接重新连接各流以维持整体一致性。实验表明,MTSS在多种模型上持续提升视频理解能力,在Video-SALMONN-2上平均误差率降低25%,在Daily-Omni推理基准上平均性能提升67%。同时缩小小规模与大规模模型间的差距,证明其更易学习。此外,无需模型改造,仅替换提示为MTSS即可在多镜头生成中带来显著提升:角色一致性提高45%,音画对齐提升56%,时间可控性提升71%。
原文摘要 · Abstract (English)
Advances in Multimodal Large Language Models (MLLMs) are transforming video captioning from a descriptive endpoint into a semantic interface for both video understanding and generation. However, the dominant paradigm still casts videos as monolithic narrative paragraphs that entangle visual, auditory, and identity information. This dense coupling not only compromises representational fidelity but also limits scalability, since even local edits can trigger global rewrites. To address this structural bottleneck, we propose Multi-Stream Scene Script (MTSS), a novel paradigm that replaces monolithic text with factorized and explicitly grounded scene descriptions. MTSS is built on two core principles: Stream Factorization, which decouples a video into complementary streams (Reference, Shot, Event, and Global), and Relational Grounding, which reconnects these isolated streams through explicit identity and temporal links to maintain holistic video consistency. Extensive experiments demonstrate that MTSS consistently enhances video understanding across various models, achieving an average reduction of 25% in the total error rate on Video-SALMONN-2 and an average performance gain of 67% on the Daily-Omni reasoning benchmark. It also narrows the performance gap between smaller and larger MLLMs, indicating a substantially more learnable caption interface. Finally, even without architectural adaptation, replacing monolithic prompts with MTSS in multi-shot video generation yields substantial human-rated improvements: a 45% boost in cross-shot identity consistency, a 56% boost in audio-visual alignment, and a 71% boost in temporal controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。