评测大模型在真实手机场景下的函数调用能力,发现参数名错误是主要失败原因。
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
- 构建真实手机助手对话数据集,模拟多轮交互中的复杂行为。
- 发现参数名错误是导致函数调用失败的主要原因,影响多种场景。
- 适合研究移动助理、大模型鲁棒性及对话系统评估的学者使用。
评估大模型在多轮人机交互中的表现面临巨大挑战,尤其源于用户行为的复杂性和多样性。本文提出 HammerBench,一个针对真实移动设备场景下大模型函数调用能力的新基准框架。该框架模拟多样化的手机助手使用场景,包含不完整指令、动态问答轨迹、意图与参数变更,以及通过代词间接使用外部信息。数据集基于主流手机应用功能和匿名用户日志构建,并结合开源模型实现低成本数据生成。HammerBench 还引入细粒度交互快照与指标,支持对每一轮对话的函数调用性能进行深入分析。通过评估多个领先大模型,我们揭示了不同参数名错误是各类交互场景中失败的关键因素,凸显大模型在移动端应用中鲁棒性提升的必要性。
原文摘要 · Abstract (English)
Evaluating the performance of LLMs in multi-turn human-agent interactions presents significant challenges, particularly due to the complexity and variability of user behavior. In this paper, we introduce HammerBench, a novel benchmark framework for assessing LLMs' function-calling capabilities in real-world, multi-turn dialogues. HammerBench simulates diverse mobile assistant use cases, incorporating imperfect instructions, dynamic question-answer trajectories, intent and argument shifts, and the indirect use of external information through pronouns. To construct this benchmark, we curate a comprehensive dataset derived from popular mobile app functionalities and anonymized user logs, complemented by a cost-effective data generation pipeline leveraging open-source models. HammerBench is further augmented with fine-grained interaction snapshots and metrics, enabling detailed evaluation of function-calling performance across individual conversational turns. We demonstrate the effectiveness of HammerBench by evaluating several leading LLMs and uncovering key performance trends. Our experiments reveal that different types of parameter name errors are a significant source of failure across different interaction scenarios, highlighting critical areas for further improvement in LLM robustness for mobile assistant applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。