新基准ETAPP评估个性化工具调用能力,兼顾用户偏好与主动行为。
Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and Proactivity
- 构建涵盖800案例的测试集,模拟多样化用户画像。
- 提出基于关键点的评估方法,减少大模型评分偏差。
- 揭示工具调用策略与微调对个性化性能的影响。
个性化工具使用对提升大语言模型在多工具交互场景中的用户偏好对齐至关重要。然而现有评测基准大多只关注文本生成个性化或直接工具调用,未同时考虑二者。本文提出新型基准ETAPP,建立沙箱环境与包含800个测试用例的综合数据集,覆盖多样用户画像。为提升评估准确性,提出基于关键点的大模型评测方法,通过人工标注每例的关键点并提供给大模型作为参考,以缓解大模型自评系统中的偏差。此外,对优秀LLMs进行评估并深入分析。研究不同工具调用策略对LLM个性化表现的影响,以及微调在此任务中的效果。验证了偏好设置与关键点评估方法的有效性。研究结果为改进个性化LLM智能体提供了重要洞见。代码已开源:https://github.com/hypasd-art/ETAPP。
原文摘要 · Abstract (English)
Personalized tool utilization is essential for aligning large language models (LLMs) with user preference in interaction scenarios with various tools. However, most of the current benchmarks primarily focus on either personalization of text generation or direct tool-utilizing, without considering both. In this work, we introduce a novel benchmark ETAPP for evaluating personalized tool invocation, establishing a sandbox environment, and a comprehensive dataset of 800 testing cases covering diverse user profiles. To improve the accuracy of our evaluation, we propose a key-point-based LLM evaluation method, mitigating biases in the LLM-as-a-judge system by manually annotating key points for each test case and providing them to LLM as the reference. Additionally, we evaluate the excellent LLMs and provide an in-depth analysis. Furthermore, we investigate the impact of different tool-invoking strategies on LLMs' personalization performance and the effects of fine-tuning in our task. The effectiveness of our preference-setting and key-point-based evaluation method is also validated. Our findings offer insights into improving personalized LLM agents. Our Code is available at https://github.com/hypasd-art/ETAPP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。