自动生成定制化大模型评估数据集,提升评测效率与可扩展性。
STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator

- 基于改进的Self-Instruct框架构建可控合成数据生成引擎
- 生成数据在大模型评测中得分比现有基准高5.7%
- 适合需要快速、低成本评测大模型的应用团队使用
随着大语言模型在各领域的广泛应用,亟需高质量的领域和语言特定评估数据集。然而,由于隐私、监管限制及人工创建耗时,这类数据集的构建极具挑战。现有自动化评估方法常依赖已有数据,存在可扩展性差、单一领域、缺乏多语言支持等问题。本文提出STELLAR-E——一个完全自动化的系统,仅需极少人工输入即可生成定制规模的高质量合成数据集,不依赖现有数据。系统分两阶段:(1) 改进TGRT Self-Instruct框架,构建可控合成数据生成引擎;(2) 设计融合统计与大模型评分的评估流水线,验证合成数据在大模型应用评测中的适用性。合成数据在LLM-as-a-judge评分上平均优于现有语言特定基准5.7%,展现出对大、小模型全面评估的可比质量。尽管真实数据对小模型仍略具挑战性,本工作建立了一个可扩展、可适配领域的评测框架,为大模型应用提供公平评估方案,显著优于人工方式,支持高效自动化质量保障循环。
原文摘要 · Abstract (English)
The increasing reliance on Large Language Models (LLMs) across diverse sectors highlights the need for robust domain-specific and language-specific evaluation datasets; however, the collection of such datasets is challenging due to privacy concerns, regulatory restrictions, and the time cost for manual creation. Existing automated benchmarking methods are often limited by relying on pre-existing data, poor scalability, single-domain focus, and lack of multilingual support. We present STELLAR-E - a fully automated system to generate high-quality synthetic datasets of custom size, using minimal human inputs without depending on existing datasets. The system is structured in two stages: (1) We modify the TGRT Self-Instruct framework to create a synthetic data engine that enables controllable, custom synthetic dataset generation, and (2) an evaluation pipeline incorporating statistical and LLM-based metrics to assess the applicability of the synthetic dataset for LLM-based application evaluations. The synthetic datasets reach an average difference of +5.7% in terms of LLM-as-a-judge scores against existing language-specific benchmarks, demonstrating comparable quality for comprehensive assessment of big and small LLMs. While real datasets remain slightly more challenging for LLMs especially for smaller models, this work establishes a scalable and domain-adaptable benchmarking framework that supports fair evaluation of LLM applications, offering a faster alternative to manual approaches and enabling high-efficiency automated quality assurance cycles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。