用视觉语言模型实现跨平台策略游戏自动操作
Yanyun-3: Enabling Cross-Platform Strategy Game Operation with Vision-Language Models
- 设计分粒度数据组织方法,区分多模态数据融合方式
- 在三平台测试中实现12.98倍BLEU-4提升与63%推理加速
- 无需定制调优即可跨平台执行核心游戏操作
跨平台策略游戏自动化面临用户界面多样与战场环境动态的挑战。现有视觉-语言模型在异构平台间泛化能力差,且对界面理解与动作执行精度不足。本文提出Yanyun-3,基于Qwen2.5-VL进行视觉推理,结合UI-TARS实现界面操作。提出新型数据组织原则——组合粒度,以区分样本内融合与样本间混合的多模态数据(静态图像、多图像序列、视频)。在三个策略游戏平台的精选数据集上,采用QLoRA微调。最优策略(M*V+S)相较全融合方式,使BLEU-4得分提升12.98倍,推理时间减少63%。Yanyun-3无需平台特化调优,即可成功完成目标选择、资源分配等核心任务。研究证明,结构化多模态数据组织显著提升视觉语言模型在具身任务中的表现,为图形用户界面自动化提供可泛化的框架,对机器人与自主系统具有更广泛意义。
原文摘要 · Abstract (English)
Cross-platform strategy game automation remains a challenge due to diverse user interfaces and dynamic battlefield environments. Existing Vision--Language Models (VLMs) struggle with generalization across heterogeneous platforms and lack precision in interface understanding and action execution. We introduce Yanyun-3, a VLM-based agent that integrates Qwen2.5-VL for visual reasoning and UI-TARS for interface execution. We propose a novel data organization principle -- combination granularity -- to distinguish intra-sample fusion and inter-sample mixing of multimodal data (static images, multi-image sequences, and videos). The model is fine-tuned using QLoRA on a curated dataset across three strategy game platforms. The optimal strategy (M*V+S) achieves a 12.98x improvement in BLEU-4 score and a 63% reduction in inference time compared to full fusion. Yanyun-3 successfully executes core tasks (e.g., target selection, resource allocation) across platforms without platform-specific tuning. Our findings demonstrate that structured multimodal data organization significantly enhances VLM performance in embodied tasks. Yanyun-3 offers a generalizable framework for GUI automation, with broader implications for robotics and autonomous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。