arXiv:2603.01055cs.AI2026-03被引 3

构建首个融合视觉的多模态常识知识图谱,提升故事生成的连贯性与情境感知。

MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning

  • 将视觉信息融入原子级常识图谱,构建90万+多模态三元组。
  • 在图像叙事任务中,生成内容更丰富、连贯且符合上下文。
  • 适合研究多模态推理与自动叙事生成的学者和开发者。

我们提出MMCOMET,首个整合物理、社会与事件性常识的多模态常识知识图谱(MMKG)。该图谱通过高效图像检索流程扩展了ATOMIC2020知识图谱,引入视觉维度,形成超过90万条多模态三元组。这一资源解决了现有MMKG在复杂推理任务(如图像描述与讲故事)中的局限性。通过标准的视觉叙事实验,我们证明该方法能生成比纯文本知识驱动更丰富、连贯且情境贴合的故事。此资源为多模态常识推理与叙事生成奠定了新基础。

原文摘要 · Abstract (English)

We present MMCOMET, the first multimodal commonsense knowledge graph (MMKG) that integrates physical, social, and eventive knowledge. MMCOMET extends the ATOMIC2020 knowledge graph to include a visual dimension, through an efficient image retrieval process, resulting in over 900K multimodal triples. This new resource addresses a major limitation of existing MMKGs in supporting complex reasoning tasks like image captioning and storytelling. Through a standard visual storytelling experiment, we show that our holistic approach enables the generation of richer, coherent, and contextually grounded stories than those produced using text-only knowledge. This resource establishes a new foundation for multimodal commonsense reasoning and narrative generation.

多模态常识推理知识图谱叙事生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。