用教别人来评估大模型,更全面且防作弊。
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
- 让大模型当老师,通过教学效果反推其能力
- 在26个主流模型上验证,结果与人工评分高度一致
- 适合想提升模型训练质量的研究者和开发者
大语言模型的发展已远超评估方法的进展。传统评测依赖特定任务指标和静态数据集,常面临公平性差、可扩展性低和数据泄露风险。本文提出 Teach2Eval,一种受费曼技巧启发的间接评估框架。不直接测试模型在预设任务上的表现,而是评估其指导弱学生模型有效完成任务的能力。通过教师生成的反馈将开放任务转化为标准化多项选择题,实现可扩展、自动化、多维度评估。该方法避免了数据泄露和记忆问题,同时捕捉到与现有基准正交的广泛认知能力。在26个领先大模型上的实验表明,其结果与现有基于人类和模型的动态排名高度一致,并为训练提供额外可解释性。
原文摘要 · Abstract (English)
Recent progress in large language models (LLMs) has outpaced the development of effective evaluation methods. Traditional benchmarks rely on task-specific metrics and static datasets, which often suffer from fairness issues, limited scalability, and contamination risks. In this paper, we introduce Teach2Eval, an indirect evaluation framework inspired by the Feynman Technique. Instead of directly testing LLMs on predefined tasks, our method evaluates a model's multiple abilities to teach weaker student models to perform tasks effectively. By converting open-ended tasks into standardized multiple-choice questions (MCQs) through teacher-generated feedback, Teach2Eval enables scalable, automated, and multi-dimensional assessment. Our approach not only avoids data leakage and memorization but also captures a broad range of cognitive abilities that are orthogonal to current benchmarks. Experimental results across 26 leading LLMs show strong alignment with existing human and model-based dynamic rankings, while offering additional interpretability for training guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。