用对抗性编码让饱和的模型评测重新有效
Resurrecting saturated LLM benchmarks with adversarial encoding
- 通过配对问题和增加选项,构造对抗性输入
- 大模型性能下降,旧基准重现未饱和状态
- 适合想重用经典评测的模型评估者
近期研究发现,微调评测问题可降低大模型的推理与记忆能力。本文在三个基准(WMDP-bio、GPQA、MMLU 变体)上测试两种修改:问题配对与增加答案选项。结果表明,对于更强大的模型,这些改动会显著降低其表现,从而打破原有性能天花板,使基准重新具备区分度。该方法可有效‘复活’已饱和的旧基准。
原文摘要 · Abstract (English)
Recent work showed that small changes in benchmark questions can reduce LLMs' reasoning and recall. We explore two such changes: pairing questions and adding more answer options, on three benchmarks: WMDP-bio, GPQA, and MMLU variants. We find that for more capable models, these predictably reduce performance, essentially heightening the performance ceiling of a benchmark and unsaturating it again. We suggest this approach can resurrect old benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。