arXiv:2603.23522cs.CLcs.AI2026-03被引 5

为每个问题定制评估标准,让大模型评测更精准。

Qworld: Question-Specific Evaluation Criteria for LLMs

  • 用递归树分解问题,生成针对性评估标准。
  • 在HealthBench上覆盖89%专家标准,产出79%新标准。
  • 适合需要细粒度评估的前沿大模型研究者。

对开放式问题评估大语言模型(LLMs)困难,因回答质量依赖于问题上下文。二元评分和静态评分标准无法捕捉这种上下文相关要求。现有方法在数据集层面定义标准或单次生成,难以探索每个问题隐含的评估空间。本文提出One-Question-One-World(Qworld),通过递归扩展树为每个问题生成特定评估标准。给定一个问题,Qworld通过层级与横向扩展,将其分解为场景、视角和细粒度二元标准,明确高质量回答需涵盖的内容。在HealthBench上,Qworld覆盖89%专家制定的标准,并生成79%经人工验证的新标准。专家评价显示,其标准在洞察力与颗粒度上优于先前方法。应用于11个前沿大模型在HealthBench与Humanity's Last Exam上的测试,揭示了长期影响、公平性、错误处理及跨学科推理等维度的能力差异,这些是粗略评分标准无法捕捉的。通过为每道题生成评估标准,Qworld实现基于问题而非任务级别的响应评估。

原文摘要 · Abstract (English)

Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Binary scores and static rubrics fail to capture these context-dependent requirements. Existing methods define criteria at the dataset level or generate them in a single pass, which limits their ability to explore the evaluation space implied by each question. We introduce One-Question-One-World (Qworld), a method that generates question-specific evaluation criteria using a recursive expansion tree. Given a question, Qworld decomposes it into scenarios, perspectives, and fine-grained binary criteria through hierarchical and horizontal expansion. The resulting criteria specify what a high-quality answer must address for that question. On HealthBench, Qworld covers 89% of expert-authored criteria and generates 79% novel criteria validated by human experts. Experts rate Qworld criteria higher in insight and granularity than those produced by prior methods. When applied to 11 frontier LLMs on HealthBench and Humanity's Last Exam, Qworld reveals capability differences in dimensions such as long-term impact, equity, error handling, and interdisciplinary reasoning that coarse rubrics do not capture. By generating evaluation criteria for each question, Qworld enables assessment of LLM responses that is tailored to the question rather than based on fixed task-level criteria.

大模型评估问答系统标准生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。