arXiv:2504.02670cs.AIcs.CL2025-04被引 9

用动态知识图谱让小模型也能高效完成复杂任务

Affordable AI Assistants with Knowledge Graph of Thoughts

  • 将任务知识构建成动态知识图谱,结合外部工具迭代优化
  • 在GAIA基准上任务成功率提升29%,成本降低36倍以上
  • 适合追求低成本、高性价比AI助手的开发者与企业

大型语言模型(LLMs)正推动AI助手在多领域执行多样化任务。然而,当前最先进的LLM驱动代理面临高昂运营成本和复杂基准(如GAIA)上成功率低的问题。为此,我们提出知识图谱思维(KGoT),一种融合LLM推理与动态构建知识图谱(KG)的新型AI助手架构。KGoT将任务相关知识提取并结构化为动态知识图谱,通过数学求解器、网络爬虫和Python脚本等外部工具持续增强。这种结构化表示使小型模型能有效解决复杂任务,同时减少偏差与噪声。例如,KGoT在GAIA基准上的任务成功率比使用GPT-4o mini的Hugging Face Agents高出29%;且采用更小模型使运营成本降低超过36倍。其他模型(如Qwen2.5-32B和Deepseek-R1-70B)及基准(如SimpleQA)也表现出类似改进。KGoT提供了一种可扩展、低成本、多功能且高性能的AI助手解决方案。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are revolutionizing the development of AI assistants capable of performing diverse tasks across domains. However, current state-of-the-art LLM-driven agents face significant challenges, including high operational costs and limited success rates on complex benchmarks like GAIA. To address these issues, we propose Knowledge Graph of Thoughts (KGoT), an innovative AI assistant architecture that integrates LLM reasoning with dynamically constructed knowledge graphs (KGs). KGoT extracts and structures task-relevant knowledge into a dynamic KG representation, iteratively enhanced through external tools such as math solvers, web crawlers, and Python scripts. Such structured representation of task-relevant knowledge enables low-cost models to solve complex tasks effectively while also minimizing bias and noise. For example, KGoT achieves a 29% improvement in task success rates on the GAIA benchmark compared to Hugging Face Agents with GPT-4o mini. Moreover, harnessing a smaller model dramatically reduces operational costs by over 36x compared to GPT-4o. Improvements for other models (e.g., Qwen2.5-32B and Deepseek-R1-70B) and benchmarks (e.g., SimpleQA) are similar. KGoT offers a scalable, affordable, versatile, and high-performing solution for AI assistants.

AI助手知识图谱低成本大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。