arXiv:2411.14054cs.CLcs.AI2024-11被引 6

评测大模型在韩语工具调用对话中的生成能力

FunctionChat-Bench: Comprehensive Evaluation of Language Models' Generative Capabilities in Korean Tool-use Dialogs

  • 按对话类型分类输出,构建四类评估维度
  • 700项测试题显示单轮调用准确率不等于多轮表现好
  • 强调对话生成能力比工具调用更重要,适合对话系统研究者

本研究探究大语言模型在工具调用对话中的生成能力。我们将模型输出分为四类:工具调用(Tool Call)、答案补全(Answer Completion)、槽位提问(Slot Question)和相关性检测(Relevance Detection),作为评估维度。我们提出 FunctionChat-Bench,包含700个评估条目及自动化评估程序。基于该基准,评估了多个支持函数调用的语言模型。结果表明,尽管模型在单轮工具调用场景中表现准确,但这并不意味着其在多轮对话环境中具备优秀的生成能力。我们认为,函数调用所需能力不仅限于生成调用指令,还需有效生成与用户互动的对话内容。

原文摘要 · Abstract (English)

This study investigates language models' generative capabilities in tool-use dialogs. We categorize the models' outputs in tool-use dialogs into four distinct types: Tool Call, Answer Completion, Slot Question, and Relevance Detection, which serve as aspects for evaluation. We introduce FunctionChat-Bench, comprising 700 evaluation items and automated assessment programs. Using this benchmark, we evaluate several language models that support function calling. Our findings indicate that while language models may exhibit high accuracy in single-turn Tool Call scenarios, this does not necessarily translate to superior generative performance in multi-turn environments. We argue that the capabilities required for function calling extend beyond generating tool call messages; they must also effectively generate conversational messages that engage the user.

工具调用对话评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。