提出评估大模型多选题表现的新协议,揭示评分指标与答案波动的强关联。
Metric assessment protocol in the context of answer fluctuation on MCQ tasks
- 基于答案波动率与原始性能分析评估指标
- 发现现有指标均与答案变化高度相关
- 新指标'最差准确率'关联性最强,适合可靠性研究
使用多选题(MCQ)已成为高效评估大语言模型能力的标准方法。可采用多种评价指标,但以往研究未对这些指标进行充分评估。同时,MCQ评估存在答案波动问题:模型在提示微小变化下会给出不同结果。本文提出一种指标评估协议,通过分析评价方法与波动率、原始性能之间的关系进行系统评估。结果表明,即使不引入额外提示变体,现有指标与答案变化之间仍存在强烈关联。新提出的指标‘最差准确率’在该协议中表现出最高关联性。
原文摘要 · Abstract (English)
Using multiple-choice questions (MCQs) has become a standard for assessing LLM capabilities efficiently. A variety of metrics can be employed for this task. However, previous research has not conducted a thorough assessment of them. At the same time, MCQ evaluation suffers from answer fluctuation: models produce different results given slight changes in prompts. We suggest a metric assessment protocol in which evaluation methodologies are analyzed through their connection with fluctuation rates, as well as original performance. Our results show that there is a strong link between existing metrics and the answer changing, even when computed without any additional prompt variants. A novel metric, worst accuracy, demonstrates the highest association on the protocol.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。