arXiv:2603.19264cs.CLcs.AI2026-03被引 1

用代理任务提升大模型测评效率,降低标注成本。

Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation

  • 将生成式问答转为伪分类任务,捕捉样本不确定性。
  • 零样本采样函数误差降低约40%,显著优于传统方法。
  • 适合需要专家标注的医疗、生物等领域模型评测。

随着预训练大语言模型(LLM)的广泛应用,医疗、生物医学等领域对特定任务测试集的需求日益增长。然而,构建新基准时标注样本的成本高昂,尤其当需专家参与时更为突出。现有主动样本选择框架对生成式问答任务支持有限,因选项动态会影响模型决策边界。本文提出生成式主动测试(Generative Active Testing, GAT),一种基于不确定性的采样框架,利用LLM作为代理以指导样本选择。通过创新的陈述适配模块,将生成式任务转化为伪分类格式,实现对未标注样本的逐样本不确定性捕捉。其零样本采样函数相比传统基线,估计误差降低约40%,提供了一种可扩展的低成本模型评估方案。

原文摘要 · Abstract (English)

With the widespread adoption of pre-trained Large Language Models (LLM), there exists a high demand for task-specific test sets to benchmark their performance in domains such as healthcare and biomedicine. However, the cost of labeling test samples while developing new benchmarks poses a significant challenge, especially when expert annotators are required. Existing frameworks for active sample selection offer limited support for generative Question Answering tasks, where option dynamics can affect model decision boundaries. In this paper, we present Generative Active Testing (GAT), an uncertainty-aware acquisition framework leveraging LLMs as surrogates for informing the sample selection process. Using a novel Statement Adaptation Module, we modify generative tasks into a pseudo-classification format, enabling the capture of sample-level uncertainties across unlabeled candidates. Our zero-shot acquisition functions reduce estimation error by ~40% compared to traditional sampling baselines, offering a scalable solution for cost-effective model benchmarking.

大模型评测主动学习生成式任务零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。