arXiv:2504.06766cs.AIcs.CL2025-04被引 3

构建家庭知识图谱评估大模型多跳个性化工具使用能力。

FamilyTool: A Multi-hop Personalized Tool Use Benchmark

  • 基于家庭关系图谱设计多跳推理任务,模拟真实个性化场景。
  • 在2-6跳未知关系下,顶尖模型准确率大幅下降,暴露泛化短板。
  • 适合研究大模型推理、自适应与动态环境交互的学者使用。

将工具学习与大型语言模型(LLMs)结合,可提升其处理复杂任务的能力。然而,现有工具学习基准未能充分覆盖真实世界中需要多跳推理和归纳知识适应的个性化场景。为此,我们提出FamilyTool,一个基于家庭知识图谱(KG)的新基准,模拟个性化、多跳工具使用场景。FamilyTool包含基础与扩展数据集,挑战模型在1至4跳(如推断亲属关系与偏好)及2至6跳的查询任务,采用归纳式KG设置,要求模型在不重新训练的前提下适应未见用户偏好与关系,这是以往方法的常见局限,影响泛化能力。我们进一步提出KGETool:一种简单的图谱增强型评估流程,系统评估模型在此类场景下的工具使用能力。实验显示,当前最先进模型在多跳任务中表现显著下滑,归纳场景更暴露出严重泛化缺陷。这些结果揭示了现有大模型在处理个性化、动态现实情境中的局限,凸显改进工具学习框架的迫切需求。FamilyTool为评估和推进大模型智能体在复杂动态环境中的推理、适应性与可扩展性提供了关键资源。代码与数据集可在https://github.com/yxzwang/FamilyTool获取。

原文摘要 · Abstract (English)

The integration of tool learning with Large Language Models (LLMs) has expanded their capabilities in handling complex tasks by leveraging external tools. However, existing benchmarks for tool learning inadequately address critical real-world personalized scenarios, particularly those requiring multi-hop reasoning and inductive knowledge adaptation in dynamic environments. To bridge this gap, we introduce FamilyTool, a novel benchmark grounded in a family-based knowledge graph (KG) that simulates personalized, multi-hop tool use scenarios. FamilyTool, including base and extended datasets, challenges LLMs with queries spanning from 1 to 4 relational hops (e.g., inferring familial connections and preferences) and 2 to 6 hops respectively, and incorporates an inductive KG setting where models must adapt to unseen user preferences and relationships without re-training, a common limitation in prior approaches that compromises generalization. We further propose KGETool: a simple KG-augmented evaluation pipeline to systematically assess LLMs' tool use ability in these settings. Experiments reveal significant performance gaps in state-of-the-art LLMs, with accuracy dropping sharply as hop complexity increases and inductive scenarios exposing severe generalization deficits. These findings underscore the limitations of current LLMs in handling personalized, evolving real-world contexts and highlight the urgent need for advancements in tool-learning frameworks. FamilyTool serves as a critical resource for evaluating and advancing LLM agents' reasoning, adaptability, and scalability in complex, dynamic environments. Code and dataset are available at \href{https://github.com/yxzwang/FamilyTool}{https://github.com/yxzwang/FamilyTool}.

大模型工具使用多跳推理知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。