arXiv:2511.21860cs.CLcs.AI2025-11被引 1

用一致性评估提升大模型选择题评分可靠性

Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices

  • 通过改写选项生成合成题,检测模型回答一致性
  • 发现高得分模型仍可能一致性差,新指标可有效降分
  • 适合关注模型真实能力而非表面分数的研究者

本文提出一致性重平衡准确率(CoRA)指标,提升大语言模型在多选题基准测试中的评分可靠性。该指标通过合成生成的、选项被修改的问题,考察模型的回答一致性,引入两个中间指标:最低一致性准确率(BMCA)和一致性指数(CI),以调整原始多选题回答准确率,更真实反映模型一致性水平。我们在多个基准和不同大模型上进行评估,不仅证实大模型即使在高多选题准确率下也可能表现出低响应一致性,还表明CoRA能有效降低不一致模型的得分。

原文摘要 · Abstract (English)

In this work we present the Consistency-Rebalanced Accuracy (CoRA) metric, improving the reliability of Large Language Model (LLM) scores computed on multiple choice (MC) benchmarks. Our metric explores the response consistency of the LLMs, taking advantage of synthetically-generated questions with altered answer choices. With two intermediate scores, i.e. Bare-Minimum-Consistency Accuracy (BMCA) and Consistency Index (CI), CoRA is computed by adjusting the multiple-choice question answering (MCQA) scores to better reflect the level of consistency of the LLM. We present evaluations in different benchmarks using diverse LLMs, and not only demonstrate that LLMs can present low response consistency even when they present high MCQA scores, but also that CoRA can successfully scale down the scores of inconsistent models.

大模型评估多选题一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。