arXiv:2603.27690cs.CV2026-03中稿 · CVPR

用多模态输入定制电影级故事生成,支持角色与镜头类型控制。

Customized Visual Storytelling with Unified Multimodal LLMs

  • 融合文本、人物图像和背景参考,实现多模态故事定制。
  • 通过参数高效提示调优,提升镜头类型多样性与电影语法还原度。
  • 新构建两个基准,评估角色一致性、画面对齐与镜头控制能力。

多模态故事定制旨在根据文本描述、参考人物图像和镜头类型生成连贯的故事流。尽管近期故事生成取得进展,多数方法仍依赖纯文本输入。少数研究引入人物身份线索(如面部识别),但缺乏更广泛的多模态条件控制。本文提出VstoryGen框架,整合描述、人物与背景参考,实现可定制的故事生成。为增强电影化多样性,我们基于电影数据采用参数高效提示调优实现镜头类型控制,使生成序列更忠实于电影语法规则。为评估该框架,我们建立两个新基准,从角色一致性、场景一致性、文本-视觉对齐及镜头类型控制角度进行评测。实验表明,VstoryGen在一致性与电影化多样性上优于现有方法。

原文摘要 · Abstract (English)

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, most approaches rely on text-only inputs. A few studies incorporate character identity cues (e.g., facial ID), but lack broader multimodal conditioning. In this work, we introduce VstoryGen, a multimodal framework that integrates descriptions with character and background references to enable customizable story generation. To enhance cinematic diversity, we introduce shot-type control via parameter-efficient prompt tuning on movie data, enabling the model to generate sequences that more faithfully reflect cinematic grammar. To evaluate our framework, we establish two new benchmarks that assess multimodal story customization from the perspectives of character and scene consistency, text-visual alignment, and shot-type control. Experiments demonstrate that VstoryGen achieves improved consistency and cinematic diversity compared to existing methods.

故事生成多模态电影化可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。