arXiv:2601.03268cs.CLcs.LG2026-01

为小模型设计新评估框架,突出其在语气修改等实用任务中的优势。

WRAVAL -- WRiting Assist eVALuation

  • 构建数据生成与提示调优结合的评估框架
  • 小模型在语气修改任务上表现接近大模型
  • 适合边缘计算和私有场景的模型选型参考

大型语言模型(LLMs)的兴起推动语言模型评估转向推理与解题任务,以衡量通用智能。小型语言模型(SLMs,参数量低于100亿)在此类指标上通常比LLMs低3-4倍。然而,我们证明这些评估无法反映SLMs在工业常见任务(如语气修改:幽默、正式、专业)中的实际效能。为此,我们提出一种专门针对非推理任务、无预设数据集场景的评估框架,融合新颖的数据生成、提示调优与基于LLM的评估方法,展示任务定制微调的潜力。本工作为实践者提供工具,有效评测SLMs与LLMs在实际应用中的表现,尤其适用于边缘计算和私有计算场景。实现代码已开源:https://github.com/amazon-science/wraval。

原文摘要 · Abstract (English)

The emergence of Large Language Models (LLMs) has shifted language model evaluation toward reasoning and problem-solving tasks as measures of general intelligence. Small Language Models (SLMs) -- defined here as models under 10B parameters -- typically score 3-4 times lower than LLMs on these metrics. However, we demonstrate that these evaluations fail to capture SLMs' effectiveness in common industrial applications, such as tone modification tasks (e.g., funny, serious, professional). We propose an evaluation framework specifically designed to highlight SLMs' capabilities in non-reasoning tasks where predefined evaluation datasets don't exist. Our framework combines novel approaches in data generation, prompt-tuning, and LLM-based evaluation to demonstrate the potential of task-specific finetuning. This work provides practitioners with tools to effectively benchmark both SLMs and LLMs for practical applications, particularly in edge and private computing scenarios. Our implementation is available at: https://github.com/amazon-science/wraval.

小模型评估语气修改边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。