arXiv:2601.01366cs.AI2026-01

针对教育类私有软件,构建多平台代理评估框架,提升跨平台任务理解与执行效率。

KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models

  • 引入知识增强双图评估框架,分解任务为子目标逐项验证。
  • 在104个教育任务上测试,显著提升私有软件场景下的代理执行效率。
  • 适合研究多模态大模型在教育智能体中的应用与评估的学者。

随着多模态大语言模型(MLMs)在自主智能体中的广泛应用,教育场景下跨平台任务执行能力受到广泛关注。然而,现有基准框架在支持教育领域跨平台任务时仍存在明显不足,尤其在处理学校专用软件(如 XiaoYa 智能助手、华师小智等)时,代理因缺乏对私有系统结构的理解而效率大幅下降。此外,当前评估方法主要依赖目标达成或轨迹匹配等粗粒度指标,难以捕捉复杂任务中代理的执行细节与效率。为此,我们提出 KGCE(Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models),一个融合知识库增强与双图评估框架的新基准平台。我们构建了包含104个教育相关任务的数据集,涵盖Windows、Android及跨平台协作任务。KGCE通过将任务分解为多个子目标并验证其完成状态,实现细粒度评估。为克服现有代理在私有领域任务中的执行瓶颈,我们开发了集成学校专用软件知识库的增强型代理系统。代码已开源:https://github.com/Kinginlife/KGCE。

原文摘要 · Abstract (English)

With the rapid adoption of multimodal large language models (MLMs) in autonomous agents, cross-platform task execution capabilities in educational settings have garnered significant attention. However, existing benchmark frameworks still exhibit notable deficiencies in supporting cross-platform tasks in educational contexts, especially when dealing with school-specific software (such as XiaoYa Intelligent Assistant, HuaShi XiaZi, etc.), where the efficiency of agents often significantly decreases due to a lack of understanding of the structural specifics of these private-domain software. Additionally, current evaluation methods heavily rely on coarse-grained metrics like goal orientation or trajectory matching, making it challenging to capture the detailed execution and efficiency of agents in complex tasks. To address these issues, we propose KGCE (Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models), a novel benchmarking platform that integrates knowledge base enhancement and a dual-graph evaluation framework. We first constructed a dataset comprising 104 education-related tasks, covering Windows, Android, and cross-platform collaborative tasks. KGCE introduces a dual-graph evaluation framework that decomposes tasks into multiple sub-goals and verifies their completion status, providing fine-grained evaluation metrics. To overcome the execution bottlenecks of existing agents in private-domain tasks, we developed an enhanced agent system incorporating a knowledge base specific to school-specific software. The code can be found at https://github.com/Kinginlife/KGCE.

教育智能体跨平台评估多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。