arXiv:2607.01312cs.CV2026-07

检测生成故事图中语义连贯性丢失,发现画面看似连贯却讲不清情节变化。

KathaTrace: Diagnosing Semantic Trajectory Collapse in Generated Visual Narratives

论文配图:KathaTrace: Diagnosing Semantic Trajectory Collapse in Generated Visual Narratives
图 1 · 摘自论文原文
  • 通过文本/图像单独与联合判断,检测场景间语义衔接是否断裂
  • 实测顶尖模型生成故事的语义断连率高达23.5%,严重偏离原意
  • 适合故事生成、影视预览等需逻辑连贯性的研究者使用

视觉叙事在分镜稿、漫画、儿童媒体和电影预演中至关重要,观众仅凭图像理解故事。近期如StoryDiffusion等生成器可产出视觉连贯序列,但视觉一致性不等于原故事的过渡意义仍可还原。现有基准评估视觉质量、内容忠实度和场景连贯性,却忽略了关键缺陷:画面看似连贯,但前后场景间的语义关联已消失。本文提出KathaTrace,一种无需依赖生成器的诊断协议,用于检测‘语义轨迹坍塌’——即理解一个场景如何过渡到下一个场景所需的意义丧失。KathaTrace在三种证据条件下评估过渡:仅文本、仅图像、文本加图像,并过滤模糊项。我们构建了KathaBench-25K数据集,包含来自《伊索寓言》《五卷书》《迦罗婆传说》等经典作品的5,000个叙事、20,000个过渡关系及28,712个可恢复性问题。定义‘语义轨迹差距’(STG)为仅文本与仅图像条件下的可恢复性之差,衡量可视化过程中损失的过渡意义。人工验证显示一致率Fleiss' kappa = 0.845。对多个前沿生成器的实验表明,其平均STG达23.5 ± 1.3。进一步设计的‘语义罗盘’行动性探测器利用KathaTrace信号实现生成后修复,提升分镜选择效果。

原文摘要 · Abstract (English)

Visual narratives are central to storyboards, comics, children's media, and film previsualization, where viewers understand stories from images alone. Recent generators such as StoryDiffusion produce coherent sequences, but visual coherence does not guarantee that source-story transition meaning remains recoverable. Existing benchmarks assess visual quality, content faithfulness, and scene coherence, but miss a critical failure mode: storyboards where scenes appear visually coherent while the semantic link between scenes disappears. We introduce KathaTrace, a generator-agnostic protocol for diagnosing semantic trajectory collapse, defined as the loss of transition meaning needed to understand how one scene follows another. KathaTrace evaluates transitions under three evidence conditions: text-only, image-only, and text-plus-image, and filters ambiguous items. We contribute KathaBench-25K, with 5,000 narratives from classical collections including Aesop, Panchatantra, and Kathasaritasagara, 20,000 transitions, and 28,712 recoverability questions. We define Semantic Trajectory Gap, or STG, as text-only minus image-only recoverability, measuring transition meaning lost during visualization. Human validation yields Fleiss' kappa = 0.845. Experiments across state-of-the-art generators show substantial STG of 23.5 +/- 1.3. Semantic Compass, an actionability probe, uses KathaTrace signals for post-generation repair and improves storyboard selection.

故事生成语义连贯评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。