构建指南遵循评测集,检验大模型在专业领域规则下的执行能力
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
- 设计多维度评测框架,覆盖规则遵守、更新适应与人类偏好对齐
- 实测显示主流大模型在领域规则遵循上仍有显著提升空间
- 适合关注大模型可靠性与专业应用落地的研究者和开发者
大型语言模型已广泛用作能执行用户指令并作出决策的自主智能体。以往研究主要聚焦于通用领域中模型的指令遵循能力,重点考察其常识知识。近年来,大模型越来越多地作为领域导向智能体使用,依赖与常识可能冲突的领域特定规则,这些规则具有种类繁多且频繁更新的特点。然而,目前缺乏全面评估大模型领域规则遵循能力的基准测试,严重制约了其性能评估与进一步发展。本文提出GuideBench,一个综合性评测基准,用于评估大模型在三大关键方面的表现:(i)对多样化规则的遵守程度,(ii)对规则变更的鲁棒性,(iii)与人类偏好的对齐度。对多种大模型的实验结果表明,其在遵循领域导向规则方面仍存在巨大改进空间。
原文摘要 · Abstract (English)
Large language models (LLMs) have been widely deployed as autonomous agents capable of following user instructions and making decisions in real-world applications. Previous studies have made notable progress in benchmarking the instruction following capabilities of LLMs in general domains, with a primary focus on their inherent commonsense knowledge. Recently, LLMs have been increasingly deployed as domain-oriented agents, which rely on domain-oriented guidelines that may conflict with their commonsense knowledge. These guidelines exhibit two key characteristics: they consist of a wide range of domain-oriented rules and are subject to frequent updates. Despite these challenges, the absence of comprehensive benchmarks for evaluating the domain-oriented guideline following capabilities of LLMs presents a significant obstacle to their effective assessment and further development. In this paper, we introduce GuideBench, a comprehensive benchmark designed to evaluate guideline following performance of LLMs. GuideBench evaluates LLMs on three critical aspects: (i) adherence to diverse rules, (ii) robustness to rule updates, and (iii) alignment with human preferences. Experimental results on a range of LLMs indicate substantial opportunities for improving their ability to follow domain-oriented guidelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。