arXiv:2410.04587cs.LGcs.AI2024-10被引 63

提出新模型Hammer,让手机端大模型更准调用外部功能。

Hammer: Robust Function-Calling for On-Device Language Models via Function Masking

  • 用函数掩码和增强数据提升模型对无关函数的敏感度。
  • 在多个基准上超越更大模型,表现稳定且领先。
  • 适合开发离线智能助手或需精准调用工具的应用。

大型语言模型在具备外部工具和API调用能力后,展现出作为自主代理的强大潜力。然而,其执行复杂任务的能力仍严重依赖于函数调用性能的提升。本文发现现有函数调用模型在不同基准间表现差异显著,常因特定命名习惯被误导。为此,我们提出Hammer,一类专为设备端函数调用设计的基础模型。Hammer采用增强数据集,提升模型对无关函数的识别能力,并引入函数掩码技术以减少干扰。实证评估表明,Hammer不仅超越更大规模模型,还在多样基准上展现强泛化能力,达到当前最优(SOTA)表现。开源内容包括用于无关性检测的专用数据集、增强泛化的微调框架及Hammer模型,为函数调用性能树立新标准。

原文摘要 · Abstract (English)

Large language models have demonstrated impressive value in performing as autonomous agents when equipped with external tools and API calls. Nonetheless, effectively harnessing their potential for executing complex tasks crucially relies on enhancements in their function calling capabilities. This paper identifies a critical gap in existing function calling models, where performance varies significantly across benchmarks, often due to being misled by specific naming conventions. To address such an issue, we introduce Hammer, a novel family of foundation models specifically engineered for on-device function calling. Hammer employs an augmented dataset that enhances models' sensitivity to irrelevant functions and incorporates function masking techniques to minimize misleading. Our empirical evaluations reveal that Hammer not only outperforms larger models but also demonstrates robust generalization across diverse benchmarks, achieving sota results. Our open source contributions include a specialized dataset for irrelevance detection, a tuning framework for enhanced generalization, and the Hammer models, establishing a new standard for function calling performance.

函数调用设备端模型AI代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。