小模型能否高效完成语法纠错与简化?实验表明仍不及大模型。
Evaluating Small Decoder-Only Language Models for Grammar Correction and Text Simplification
- 直接测试小型解码器模型在两项任务中的表现
- 小模型在保留语义和避免幻觉上表现不佳,性能低于主流大模型
- 适合关注轻量级模型部署的读者,但需了解其局限性
大型语言模型因在文本生成与重写等任务中表现优异而备受关注,但其庞大的规模和高计算成本使其在许多场景下难以访问、部署与保障安全。本文探究小型解码器语言模型(SLMs)是否可作为语法纠错与文本简化任务的高效替代方案。实验在JFLEG和ASSET数据集上,对未微调及微调后的小型模型进行评估,使用标准指标衡量表现。结果表明,尽管部分行为可被学习,但小模型整体性能仍低于强基线与当前大型语言模型。此外,小模型在保持原意和防止幻觉方面存在明显困难。这些发现说明,尽管具备效率优势,现有小型模型在重写任务上尚未达到现代大型模型的竞争力,未来训练方法需进一步突破以缩小性能差距。
原文摘要 · Abstract (English)
Large language models have become extremely popular recently due to their ability to achieve strong performance on a variety of tasks, such as text generation and rewriting, but their size and computation cost make them difficult to access, deploy, and secure in many settings. This paper investigates whether small, decoder-only language models can provide an efficient alternative for the tasks of grammar correction and text simplification. The experiments in this paper focus on testing small language models out of the box, fine-tuned, and run sequentially on the JFLEG and ASSET datasets using established metrics. The results show that while SLMs may learn certain behaviors well, their performance remains below strong baselines and current LLMs. The results also show that SLMs struggle with retaining meaning and hallucinations. These findings suggest that despite their efficiency advantages, current SLMs are not yet competitive enough with modern LLMs for rewriting, and further advances in training are required for SLMs to close the performance gap between them and today's LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。