用自然语言测试项精准评估大模型表现,提升评测可靠性和开发效率。
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
- 将模型输出拆解为可测试的自然语言规则,实现细粒度评估
- 在FLASK和BigGenBench上达到顶尖水平,评分一致性显著提高
- 适合需要高精度评测的LLM研发团队和评估工具开发者
随着大语言模型融入关键工作流,其行为评估仍是核心挑战——人工评价成本高且易有噪音,而自动指标仅提供粗略、难以解读的信号。本文提出自然语言单元测试范式,将响应质量分解为明确、可测试的标准,并引入统一评分模型LMUnit,融合偏好判断、直接评分与自然语言推理三类信号进行多目标训练。通过受控的人工实验,验证该范式显著提升标注者间一致性,支持更高效的LLM开发流程。LMUnit在FLASK、BigGenBench等评估基准上达到当前最优性能,在RewardBench上表现具竞争力。结果同时验证了新范式与模型的有效性,为语言模型评估与开发提供了新路径。
原文摘要 · Abstract (English)
As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge -- human evaluation is costly and noisy, while automated metrics provide only coarse, difficult-to-interpret signals. We introduce natural language unit tests, a paradigm that decomposes response quality into explicit, testable criteria, along with a unified scoring model, LMUnit, which combines multi-objective training across preferences, direct ratings, and natural language rationales. Through controlled human studies, we show this paradigm significantly improves inter-annotator agreement and enables more effective LLM development workflows. LMUnit achieves state-of-the-art performance on evaluation benchmarks (FLASK, BigGenBench) and competitive results on RewardBench. These results validate both our proposed paradigm and scoring model, suggesting a promising path forward for language model evaluation and development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。