为俄语大模型设计可解释评估框架,用模型自评替代人工打分。
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
- 用模型互评+理由说明,实现透明化评估。
- 覆盖35类任务共2100个专业级提示,含难易分级。
- 适合开发俄语生成模型的研究者与评测团队使用。
我们提出POLLUX,一个面向俄语大语言模型生成能力的开源综合评估基准。核心贡献是新型评估方法:针对每类任务定义详细标准,建立评分协议,让模型对输出进行评价并给出理由,实现超越传统耗时的人工对比的可解释评估。POLLUX包含35种细粒度任务类型,覆盖代码生成、创意写作和实用助手等多样场景,共2100个由专家手工设计并专业撰写的提示,每项任务标注了难度等级(简单/中等/困难),数据集完全从零构建。同时发布一组用于细致评估生成结果的LLM-as-a-Judge(7B和32B)评估器。该方法提供可扩展、可解释的评估与标注工具,有效替代成本高且精度有限的人工判断。
原文摘要 · Abstract (English)
We introduce POLLUX, a comprehensive open-source benchmark designed to evaluate the generative capabilities of large language models (LLMs) in Russian. Our main contribution is a novel evaluation methodology that enhances the interpretability of LLM assessment. For each task type, we define a set of detailed criteria and develop a scoring protocol where models evaluate responses and provide justifications for their ratings. This enables transparent, criteria-driven evaluation beyond traditional resource-consuming, side-by-side human comparisons. POLLUX includes a detailed, fine-grained taxonomy of 35 task types covering diverse generative domains such as code generation, creative writing, and practical assistant use cases, totaling 2,100 manually crafted and professionally authored prompts. Each task is categorized by difficulty (easy/medium/hard), with experts constructing the dataset entirely from scratch. We also release a family of LLM-as-a-Judge (7B and 32B) evaluators trained for nuanced assessment of generative outputs. This approach provides scalable, interpretable evaluation and annotation tools for model development, effectively replacing costly and less precise human judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。