arXiv:2608.24228cs.AI2026-08

评估大模型生成多个结果时的多样性与有效性,发现传统方法会遗漏关键差异。

Evaluating Multiple LLM Generations with Validated Task Coverage

  • 构建五领域基准VTC-Bench,自动验证输出质量与任务相关差异性
  • 提出有效任务覆盖率(VTC)指标,衡量k次生成中不同有用结果的数量
  • 揭示单次输出评估忽略的模型行为差异,适合多候选生成场景研究者

许多大模型应用在提供多个候选输出以供比较、验证或组合时最具价值。然而,主流评估仍聚焦于单一输出,或将多个样本简化为一个成功或选择的答案,可能忽略输出是否包含多个真正不同的有用结果。本文提出VTC-Bench——一个涵盖五个领域的基准,配套核心评估指标有效任务覆盖率(VTC)。该基准基于真实数据任务构建,可自动、可复现地验证输出质量与任务相关的差异性,无需依赖模型判别器。VTC衡量在k次生成中获得的不同有用结果数量。在多个模型和推理设置下,该基准得出与传统评估截然不同的结论:从单次生成质量看表现最佳的配置,并不一定具有最佳覆盖能力;简单的输出变化度量也无法可靠反映任务相关的覆盖情况。结果表明,有限候选集本身可作为评估对象,揭示出传统逐输出评估所无法察觉的模型行为差异。

原文摘要 · Abstract (English)

Many LLM applications are most useful when they provide several candidate outputs for comparison, validation, or combination. Predominant evaluation settings, however, still focus on individual outputs or reduce multiple samples to a single success or selected answer. This can miss whether the outputs include several genuinely different useful results. We introduce VTC-Bench, a five-domain benchmark for this setting, together with Validated Task Coverage (VTC) as its core evaluation quantity. The benchmark is built from carefully selected real-data tasks where both output quality and task-relevant distinctness can be checked automatically and reproducibly, without model-based judges. VTC measures how many distinct useful results are obtained within $k$ attempts. Across multiple models and inference settings, the benchmark leads to different conclusions from conventional evaluation: configurations that look strongest from single-draw quality are not necessarily those with the best coverage, and simple measures of output variation do not reliably recover task-relevant coverage. These results show that finite candidate sets can be evaluated directly as objects of interest, revealing differences in model behavior that are not apparent from conventional per-output evaluation.

大模型评估多候选生成任务覆盖基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。