arXiv:2505.10292cs.CVcs.CL2025-05被引 3

构建可追踪角色与动作的影视故事数据集,提升生成故事的一致性与创意。

StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation

  • 通过视觉相似性与人脸识别实现跨帧实体重识别。
  • 引入思维链推理,显式建模多帧间关系,减少4.06至3.56次/故事的指代幻觉。
  • 适合做视觉叙事、故事生成与具身认知研究的团队使用。

视觉叙事系统常因角色身份不连贯和动作主体错配导致指代幻觉。我们提出StoryReasoning数据集,包含从52,016张电影图像中衍生的4,178个故事,提供结构化场景分析与具身故事。每个故事在多帧间保持角色与物体一致性,并通过结构化表格显式建模跨帧关系。方法包括基于视觉相似性和人脸识别的跨帧重识别、用于显式叙事建模的思维链推理,以及将文本元素与多帧视觉实体对齐的具身方案。通过微调Qwen2.5-VL 7B构建Qwen Storyteller模型,实现端到端目标检测、重识别与关键点检测,维持故事中对象引用一致。评估显示,相比未微调模型,平均每故事幻觉数由4.06降至3.56(下降12.3%),创造力由2.58提升至3.38(提升31.0%)。

原文摘要 · Abstract (English)

Visual storytelling systems struggle to maintain character identity across frames and link actions to appropriate subjects, frequently leading to referential hallucinations. These issues can be addressed through grounding of characters, objects, and other entities on the visual elements. We propose StoryReasoning, a dataset containing 4,178 stories derived from 52,016 movie images, with both structured scene analyses and grounded stories. Each story maintains character and object consistency across frames while explicitly modeling multi-frame relationships through structured tabular representations. Our approach features cross-frame object re-identification using visual similarity and face recognition, chain-of-thought reasoning for explicit narrative modeling, and a grounding scheme that links textual elements to visual entities across multiple frames. We establish baseline performance by fine-tuning Qwen2.5-VL 7B, creating Qwen Storyteller, which performs end-to-end object detection, re-identification, and landmark detection while maintaining consistent object references throughout the story. Evaluation demonstrates a reduction from 4.06 to 3.56 (-12.3%) hallucinations on average per story and an improvement in creativity from 2.58 to 3.38 (+31.0%) when compared to a non-fine-tuned model.

视觉叙事思维链具身生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。