arXiv:2602.21829cs.CVcs.AI2026-02

构建电影剧本与字幕对齐数据集,提升视觉故事生成的对话归属准确性

StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles

  • 通过最长公共子序列匹配对齐电影剧本与字幕时间戳
  • 微调后模型在字幕对齐任务中胜率达89.9%(对比基线)
  • 适合关注多模态对话生成与角色关系建模的研究者

视觉叙事模型虽能正确关联图像中的实体,仍可能虚构语义关系,导致对话归属错误、角色互动或情绪状态失真。我们提出StoryMovie,一个包含1,757个故事的数据集,通过最长公共子序列(LCS)匹配将视觉故事与电影剧本及字幕对齐。对齐流程将剧本对话与字幕时间戳同步,实现通过角色名从剧本到字幕时间位置的映射,从而支持精准对话归属。基于此对齐内容,我们生成兼具视觉标记与真实角色名、对话和关系动态的故事。在此数据集上微调Qwen Storyteller3,延续视觉接地与实体重识别工作。以DeepSeek V3为裁判评估显示,Storyteller3在字幕对齐任务中相较base Qwen2.5-VL 7B获得89.9%胜率;相比未使用剧本接地的Storyteller,分别达到48.5%与38.0%,证实语义对齐可显著提升对话归属能力。

原文摘要 · Abstract (English)

Visual storytelling models that correctly ground entities in images may still hallucinate semantic relationships, generating incorrect dialogue attribution, character interactions, or emotional states. We introduce StoryMovie, a dataset of 1,757 stories aligned with movie scripts and subtitles through LCS matching. Our alignment pipeline synchronizes screenplay dialogue with subtitle timestamps, enabling dialogue attribution by linking character names from scripts to temporal positions from subtitles. Using this aligned content, we generate stories that maintain visual grounding tags while incorporating authentic character names, dialogue, and relationship dynamics. We fine-tune Qwen Storyteller3 on this dataset, building on prior work in visual grounding and entity re-identification. Evaluation using DeepSeek V3 as judge shows that Storyteller3 achieves an 89.9% win rate against base Qwen2.5-VL 7B on subtitle alignment. Compared to Storyteller, trained without script grounding, Storyteller3 achieves 48.5% versus 38.0%, confirming that semantic alignment progressively improves dialogue attribution beyond visual grounding alone.

视觉故事对话生成数据集多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。