arXiv:2502.02028cs.CLcs.AI2025-02被引 5

小模型也能生成高质量菜谱,关键在专用评估与安全替换。

Fine-tuning Language Models for Recipe Generation: A Comparative Analysis and Benchmark Study

  • 用小模型微调生成菜谱,设计领域特化评估指标。
  • SmolLM-360M与1.7B性能相当,大模型未必更优。
  • 引入过敏原替换机制,兼顾安全性与实用性。

本研究通过微调多种极小型语言模型,探索菜谱生成任务,重点构建稳健的评估指标,并对比不同模型在开放式菜谱生成任务中的表现。实验涵盖T5-small、SmolLM-135M、Phi-2等模型架构,结合传统NLP指标与自定义领域指标。提出的新型评估框架引入菜谱内容质量评价指标,并设计过敏原替代方法。结果表明,尽管大模型在通用指标上表现更好,但在领域特定指标下,模型大小与菜谱质量的关系更为复杂:SmolLM-360M与SmolLM-1.7B在微调后性能相近;而参数量更大的Phi-2在菜谱生成中表现受限。该综合评估框架与过敏原替换系统为需要领域专长与安全考量的自然语言生成任务提供了重要参考。

原文摘要 · Abstract (English)

This research presents an exploration and study of the recipe generation task by fine-tuning various very small language models, with a focus on developing robust evaluation metrics and comparing across different language models the open-ended task of recipe generation. This study presents extensive experiments with multiple model architectures, ranging from T5-small (Raffel et al., 2023) and SmolLM-135M(Allal et al., 2024) to Phi-2 (Research, 2023), implementing both traditional NLP metrics and custom domain-specific evaluation metrics. Our novel evaluation framework incorporates recipe-specific metrics for assessing content quality and introduces approaches to allergen substitution. The results indicate that, while larger models generally perform better on standard metrics, the relationship between model size and recipe quality is more nuanced when considering domain-specific metrics. SmolLM-360M and SmolLM-1.7B demonstrate comparable performance despite their size difference before and after fine-tuning, while fine-tuning Phi-2 shows notable limitations in recipe generation despite its larger parameter count. The comprehensive evaluation framework and allergen substitution systems provide valuable insights for future work in recipe generation and broader NLG tasks that require domain expertise and safety considerations.

菜谱生成小模型评估框架安全生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。