arXiv:2505.24189cs.LGcs.AI2025-05中稿 · Workshop on Struct…被引 7

微调小模型生成低代码工作流,效果比提示大模型好10%。

Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows

  • 微调小型语言模型生成结构化输出
  • 相比提示大模型,质量提升平均10%
  • 适合需要高精度结构化输出的开发者

大型语言模型(如GPT-4o)通过恰当提示可处理多种复杂任务。随着令牌成本降低,微调小型语言模型(SLMs)在实际应用中的优势——更快推理、更低开销——可能不再明显。本文针对需生成结构化输出的领域特定任务,比较了微调SLM与提示LLM在生成低代码工作流(JSON格式)上的表现。结果表明,尽管优质提示可获得合理结果,但微调使生成质量平均提升10%。我们还进行了系统性错误分析,揭示模型在结构一致性与语义准确性方面的局限。

原文摘要 · Abstract (English)

Large Language Models (LLMs) such as GPT-4o can handle a wide range of complex tasks with the right prompt. As per token costs are reduced, the advantages of fine-tuning Small Language Models (SLMs) for real-world applications -- faster inference, lower costs -- may no longer be clear. In this work, we present evidence that, for domain-specific tasks that require structured outputs, SLMs still have a quality advantage. We compare fine-tuning an SLM against prompting LLMs on the task of generating low-code workflows in JSON form. We observe that while a good prompt can yield reasonable results, fine-tuning improves quality by 10% on average. We also perform systematic error analysis to reveal model limitations.

低代码小模型结构化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。