8B小模型作为裁判,评测性能超越GPT-4o-mini
Atla Selene Mini: A General Purpose Evaluation Model
- 用合成批评数据增强公开数据,提升评测质量
- 在11个分布外任务中表现最优,RewardBench上8B模型第一
- 零样本对齐专家评分,适合真实场景部署
我们提出Atla Selene Mini,一种前沿的小型语言模型作为裁判(SLMJ)。该模型在11个分布外基准测试中整体表现优于现有最佳SLMJ和GPT-4o-mini,涵盖绝对评分、分类与成对偏好任务。它是RewardBench上得分最高的8B生成模型,超越GPT-4o等强基线与专用裁判模型。通过系统性数据筛选与合成批评数据增强,结合直接偏好优化(DPO)与监督微调(SFT)联合训练,构建出高度可提示的评估器。在金融与医疗领域数据集上,其零样本一致性显著提升,且对提示格式变化具有鲁棒性。初步结果显示,它在实时社区驱动的Judge Arena中排名第一。模型权重已开源至HuggingFace与Ollama,推动社区广泛应用。
原文摘要 · Abstract (English)
We introduce Atla Selene Mini, a state-of-the-art small language model-as-a-judge (SLMJ). Selene Mini is a general-purpose evaluator that outperforms the best SLMJs and GPT-4o-mini on overall performance across 11 out-of-distribution benchmarks, spanning absolute scoring, classification, and pairwise preference tasks. It is the highest-scoring 8B generative model on RewardBench, surpassing strong baselines like GPT-4o and specialized judges. To achieve this, we develop a principled data curation strategy that augments public datasets with synthetically generated critiques and ensures high quality through filtering and dataset ablations. We train our model on a combined direct preference optimization (DPO) and supervised fine-tuning (SFT) loss, and produce a highly promptable evaluator that excels in real-world scenarios. Selene Mini shows dramatically improved zero-shot agreement with human expert evaluations on financial and medical industry datasets. It is also robust to variations in prompt format. Preliminary results indicate that Selene Mini is the top-ranking evaluator in a live, community-driven Judge Arena. We release the model weights on HuggingFace (https://hf.co/AtlaAI/Selene-1-Mini-Llama-3.1-8B) and Ollama to encourage widespread community adoption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。