对比人类与机器在多模态叙事中指代表达的差异,揭示模型对上下文范围的把握不足。
Coreference as an indicator of context scope in multimodal narrative
- 通过指代模式分析人类与机器生成文本的差异
- 机器在混合指代追踪上表现较差,即使生成质量看似提升
- 适合关注多模态生成上下文一致性研究者阅读
我们发现大型多模态语言模型在视觉叙事任务中,其指代表达分布与人类存在显著差异。本文提出一系列指标,量化人类与机器生成文本中指代模式的特征。人类写作时能保持跨文本与图像的一致性,以高度多样化的方式交错引用不同实体;而机器虽在生成质量上表现改善,却难以有效追踪混合指代。相关材料、度量方法与代码已公开于 https://github.com/GU-CLASP/coreference-context-scope。
原文摘要 · Abstract (English)
We demonstrate that large multimodal language models differ substantially from humans in the distribution of coreferential expressions in a visual storytelling task. We introduce a number of metrics to quantify the characteristics of coreferential patterns in both human- and machine-written texts. Humans distribute coreferential expressions in a way that maintains consistency across texts and images, interleaving references to different entities in a highly varied way. Machines are less able to track mixed references, despite achieving perceived improvements in generation quality. Materials, metrics, and code for our study are available at https://github.com/GU-CLASP/coreference-context-scope.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。