arXiv:2508.14718cs.CL2025-08

用定制分词提升菜谱生成质量,大模型表现远超传统方法

The Digital Sous Chef -- A Comparative Study on Fine-Tuning Language Models for Recipe Generation

  • 采用定制分词策略,加入23种常见分数和结构标记
  • 大模型在BERTScore上比最强基线高20%以上,困惑度降69.8%
  • 适合研究食谱生成、自然语言生成与领域适配的读者

我们建立了文本菜谱生成的严谨基准,对基于GPT-2大模型(774M)与小模型(124M)及传统LSTM/RNN基线在RecipeDB的5类菜系数据集上的表现进行了全面对比。关键贡献是提出一种针对性分词策略,将23个常见分数项和自定义结构标记加入词汇表,有效保留菜谱中的精确数值与结构信息。使用七种自动评估指标(包括BLEU-4、METEOR、ROUGE-L、BERTScore等)衡量流畅性、连贯性、语义相关性和多样性。实验表明,大模型在BERTScore(F1)上达到0.92,相比最佳循环基线(0.72)提升超过20%,困惑度降低69.8%。研究指出事实准确性仍存挑战,并为未来融合真实约束与多模态输入的高级菜谱生成研究奠定基础。

原文摘要 · Abstract (English)

We established a rigorous benchmark for text-based recipe generation, a fundamental task in natural language generation. We present a comprehensive comparative study contrasting a fine-tuned GPT-2 large (774M) model against the GPT-2 small (124M) model and traditional LSTM/RNN baselines on the 5-cuisine corpus from RecipeDB. Our key contribution is a targeted tokenization strategy that augments the vocabulary with 23 common fraction tokens and custom structural markers. This approach addresses a critical limitation of generic tokenizers by preserving essential recipe structures and precise numerical quantities, thereby enhancing domain specificity. Performance is evaluated using a comprehensive suite of seven automatic metrics spanning fluency (BLEU-4, METEOR), coherence (ROUGE-L), semantic relevance (BERTScore), and diversity. Our experiments show that the large transformer-based approach yields a >20% relative improvement in BERTScore (F1) (0.92 vs 0.72) over the best recurrent baseline, while reducing perplexity by 69.8%. We conclude with a discussion of remaining challenges, particularly regarding factual accuracy, and outline how this foundational study paves the way for integrating real-world constraints and multi-modal inputs in advanced recipe generation research.

菜谱生成大模型分词优化自然语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。