arXiv:2508.18722cs.AI2025-08EMNLP被引 16

用少量数据训练出高效游戏智能体,降低开发成本

VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft

  • 构建跨模态知识图谱融合视觉与文本信息
  • 仅需数百样本即可达成顶尖性能
  • 适合想低成本部署游戏智能体的研究者

大型语言模型在虚拟开放世界任务中表现优异,但受限于缺乏领域知识。传统微调方法依赖海量领域数据,开发成本高昂。本文提出VistaWise,一种低成本智能体框架,通过整合跨模态领域知识,并微调专用目标检测模型进行视觉分析,将领域数据需求从百万级降至数百级。该框架将视觉信息与文本依赖关系融入跨模态知识图谱(KG),实现对多模态环境的全面理解。同时,引入基于检索的池化策略从KG中提取任务相关知识,并配备桌面级技能库,支持通过鼠标键盘直接操控Minecraft客户端。实验表明,VistaWise在多种开放世界任务中达到当前最优表现,验证了其在降低开发成本的同时提升智能体性能的有效性。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown significant promise in embodied decision-making tasks within virtual open-world environments. Nonetheless, their performance is hindered by the absence of domain-specific knowledge. Methods that finetune on large-scale domain-specific data entail prohibitive development costs. This paper introduces VistaWise, a cost-effective agent framework that integrates cross-modal domain knowledge and finetunes a dedicated object detection model for visual analysis. It reduces the requirement for domain-specific training data from millions of samples to a few hundred. VistaWise integrates visual information and textual dependencies into a cross-modal knowledge graph (KG), enabling a comprehensive and accurate understanding of multimodal environments. We also equip the agent with a retrieval-based pooling strategy to extract task-related information from the KG, and a desktop-level skill library to support direct operation of the Minecraft desktop client via mouse and keyboard inputs. Experimental results demonstrate that VistaWise achieves state-of-the-art performance across various open-world tasks, highlighting its effectiveness in reducing development costs while enhancing agent performance.

游戏智能体知识图谱低成本训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。