arXiv:2410.06089cs.CLcs.AI2024-10EMNLP被引 1

用树形结构量化复杂指令重要性,提升大模型评测准确性

TOWER: Tree Organized Weighting for Evaluating Complex Instructions

  • 基于树形结构建模指令各部分重要性,模拟人类判断逻辑
  • 人工标注与树结构一致性接近人与人之间的一致性水平
  • 开源了InFoBench的树状标注数据与评估代码,便于复现

评估大语言模型遵循复杂人类编写指令的能力对实际应用部署至关重要。尽管像Chatbot Arena这类基准使用人工评判,但成本高、耗时长;而使用大模型作为评判者的方法(如AlpacaEval、MT Bench、WildBench和InFoBench)虽有改进,仍未能体现复杂指令中不同部分的重要性差异。为此,我们提出新型评估指标TOWER,将人工判断的重要性纳入复杂指令遵循评估。实验表明,人工标注者对树形结构表示的指令理解,与彼此间的一致性几乎相同。我们发布了InFoBench数据集的树状标注及配套评估代码,以推动后续研究。

原文摘要 · Abstract (English)

Evaluating the ability of large language models (LLMs) to follow complex human-written instructions is essential for their deployment in real-world applications. While benchmarks like Chatbot Arena use human judges to assess model performance, they are resource-intensive and time-consuming. Alternative methods using LLMs as judges, such as AlpacaEval, MT Bench, WildBench, and InFoBench offer improvements but still do not capture that certain complex instruction aspects are more important than others to follow. To address this gap, we propose a novel evaluation metric, \textsc{TOWER}, that incorporates human-judged importance into the assessment of complex instruction following. We show that human annotators agree with tree-based representations of these complex instructions nearly as much as they agree with other human annotators. We release tree-based annotations of the InFoBench dataset and the corresponding evaluation code to facilitate future research.

大模型评测指令遵循树形结构自动评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。