用近似奈曼分配提升大模型生成任务的高效评测
Active Testing of Large Language Models via Approximate Neyman Allocation

- 用替代模型的语义熵分层,再做近似奈曼分配采样
- 相比均匀采样降低28%误差,平均节省22.9%标注成本
- 特别适合需要专家标注的大模型生成任务评测
大语言模型从预训练到测试时扩展均需可靠评估,评估成本随模型规模和任务复杂度快速上升。主动测试通过从评估池中选取少量有信息量的样本,缓解这一瓶颈。现有方法多针对分类任务,在生成任务上表现不佳。本文提出一种专为生成任务设计的主动测试算法:利用替代模型的语义熵对评估池进行分层,并基于这些替代模型提取的信号执行近似奈曼分配。在多个语言与多模态基准及不同替代-目标模型组合下,该方法显著优于基线,接近最优的Oracle-Neyman策略,相较均匀采样最高降低28%均方误差,平均节省22.9%评估预算。
原文摘要 · Abstract (English)
Large language models (LLMs) require reliable evaluation from pre-training to test-time scaling, making evaluation a recurring rather than one-off cost. As model scales grow and target tasks increasingly demand expert annotators, both the compute and labeling costs needed for each evaluation rise rapidly. Active testing aims to alleviate this bottleneck by approximating the evaluation result from a small but informative subset of the evaluation pool. However, existing approaches primarily target classification and break down on generative tasks. We introduce a novel active testing algorithm tailored to generative tasks. Our method leverages semantic entropy from surrogate models to stratify the evaluation pool and then conducts approximate Neyman allocation based on signals extracted from these surrogates. Across multiple language and multimodal benchmarks and a range of surrogate-target model pairs, our method significantly improves on baselines and closely tracks Oracle-Neyman, delivering up to 28% MSE reduction over Uniform Sampling and an average of 22.9% budget savings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。