arXiv:2506.17589cs.AI2025-06ICCV被引 5

用游戏知识图谱让多模态模型学会自主查知识,攻克陌生任务

Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown

  • 构建游戏多模态知识图谱,关联视觉、文本与复杂实体关系
  • 无需额外训练,多智能体检索器使模型自主查找并推理知识
  • 在怪物猎人任务中显著提升模型表现,适合研究多模态推理者

知识的价值不仅在于积累,更在于有效利用以应对未知。尽管近期多模态大语言模型(MLLMs)展现出出色的多模态能力,但在罕见领域任务中常因缺乏相关知识而失败。为此,我们以视觉游戏认知为测试场景,选取《怪物猎人:世界》构建多模态知识图谱(MH-MMKG),融合多模态信息与复杂实体关系,并设计一系列挑战性查询评估模型的复杂知识检索与推理能力。此外,我们提出一种无需额外训练的多智能体检索器,使模型能自主搜索相关知识。实验表明,该方法显著提升了MLLMs性能,为多模态知识增强推理提供了新视角,并为后续研究奠定基础。

原文摘要 · Abstract (English)

The real value of knowledge lies not just in its accumulation, but in its potential to be harnessed effectively to conquer the unknown. Although recent multimodal large language models (MLLMs) exhibit impressing multimodal capabilities, they often fail in rarely encountered domain-specific tasks due to limited relevant knowledge. To explore this, we adopt visual game cognition as a testbed and select Monster Hunter: World as the target to construct a multimodal knowledge graph (MH-MMKG), which incorporates multi-modalities and intricate entity relations. We also design a series of challenging queries based on MH-MMKG to evaluate the models' ability for complex knowledge retrieval and reasoning. Furthermore, we propose a multi-agent retriever that enables a model to autonomously search relevant knowledge without additional training. Experimental results show that our approach significantly enhances the performance of MLLMs, providing a new perspective on multimodal knowledge-augmented reasoning and laying a solid foundation for future research.

多模态知识图谱推理游戏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。