对比人类与视觉语言模型在图文故事中的叙事连贯性差异。
Humans vs Vision-Language Models: A Unified Measure of Narrative Coherence
- 用多个维度指标衡量图文故事的叙事连贯性。
- 模型与人类在连贯性模式上存在系统性差异。
- 适合关注生成文本逻辑结构的研究者参考。
我们通过比较人类撰写的故事与视觉语言模型(VLMs)在Visual Writing Prompts数据集上的生成结果,研究了图文故事中的叙事连贯性。采用一组涵盖指代消解、话语关系类型、话题连续性、角色持续性以及多模态角色定位等维度的指标,计算叙事连贯性得分。结果发现,VLM生成的故事整体连贯性模式与人类存在系统性差异,尽管个别指标差异细微,但综合来看更为显著。总体表明,尽管模型输出具有类人表面流畅性,但在跨图文故事的语篇组织方式上仍与人类有本质不同。代码已开源:https://github.com/GU-CLASP/coherence-driven-humans。
原文摘要 · Abstract (English)
We study narrative coherence in visually grounded stories by comparing human-written narratives with those generated by vision-language models (VLMs) on the Visual Writing Prompts corpus. Using a set of metrics that capture different aspects of narrative coherence, including coreference, discourse relation types, topic continuity, character persistence, and multimodal character grounding, we compute a narrative coherence score. We find that VLMs show broadly similar coherence profiles that differ systematically from those of humans. In addition, differences for individual measures are often subtle, but they become clearer when considered jointly. Overall, our results indicate that, despite human-like surface fluency, model narratives exhibit systematic differences from those of humans in how they organise discourse across a visually grounded story. Our code is available at https://github.com/GU-CLASP/coherence-driven-humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。