arXiv:2601.22699cs.CL2026-01

让模型自己选最适合的答题格式,能更真实反映其能力。

Models Know Models Best: Evaluation via Model-Preferred Formats

  • 用模型自身偏好信号训练轻量分类器,自动选择最优评测格式。
  • 在零样本测试中,多个基准上准确率显著提升,表现更稳定。
  • 适合想真实评估模型推理与知识能力的研究者使用。

大型语言模型在多选题任务中的表现,因采用符号式或填空式评估格式而有显著差异。这种差异可系统归因于任务特性:自然语言续写受益于概率评分,而明确比较则更适合符号选择。该趋势在多种基于解码器的LLM中一致存在,表明其具有模型无关性。为解决此不一致性,本文提出一种动态格式对齐策略,利用轻量级分类器,基于模型内部生成的偏好信号自动判断每道题的最优格式。相比人工设计的启发式规则(常导致性能下降),该方法通过模型自身信号确定最佳评测方式,在推理与知识类基准上实现显著且稳定的零样本准确率提升,更真实揭示模型潜在能力。

原文摘要 · Abstract (English)

Performance of Large Language Models (LLMs) on multiple-choice tasks differs markedly between symbol-based and cloze-style evaluation formats. The observed discrepancies are systematically attributable to task characteristics: natural language continuation benefits from likelihood scoring, whereas explicit comparison is better suited to symbol-based selection. These trends are consistent across various decoder-based LLMs, indicating model-agnostic effects. To address these inconsistencies, a dynamic format-alignment strategy is introduced that employs a lightweight classifier trained on latent model-preference signals. In contrast to human-designed heuristics, which often degrade performance, this approach uses model-generated signals to determine the optimal format for each problem instance. The proposed method achieves substantial and consistent improvements in zero-shot accuracy across reasoning and knowledge benchmarks, better revealing the models' latent capabilities.

大模型评估零样本评测格式模型偏好

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。