arXiv:2602.15862cs.CLcs.AI2026-02被引 3

提升菜谱生成的语义准确性,让步骤和食材更合理。

Enhancing Action and Ingredient Modeling for Semantically Grounded Recipe Generation

  • 用两阶段训练增强动作与食材预测能力
  • 在Recipe1M上实现最佳语义保真度
  • 适合需要真实可靠菜谱的应用场景

多模态大模型虽能从食物图像生成菜谱,但常出现语义错误的动作或食材,尽管词法指标(如BLEU、ROUGE)表现良好。为此,我们提出一种语义根基框架,将动作与食材预测作为生成指令的内部上下文。采用两阶段流程:第一阶段通过监督微调(SFT)利用动作推理数据集和食材语料库建立基础准确率;第二阶段通过频率感知奖励进行强化微调,提升长尾动作预测与食材泛化能力。此外,引入语义置信度评分与修正(SCSR)模块,进一步过滤并纠正预测结果。在Recipe1M上的实验表明,该方法达到当前最优性能,显著提升语义一致性。

原文摘要 · Abstract (English)

Recent advances in Multimodal Large Language Models (MLMMs) have enabled recipe generation from food images, yet outputs often contain semantically incorrect actions or ingredients despite high lexical scores (e.g., BLEU, ROUGE). To address this gap, we propose a semantically grounded framework that predicts and validates actions and ingredients as internal context for instruction generation. Our two-stage pipeline combines supervised fine-tuning (SFT) with reinforcement fine-tuning (RFT): SFT builds foundational accuracy using an Action-Reasoning dataset and ingredient corpus, while RFT employs frequency-aware rewards to improve long-tail action prediction and ingredient generalization. A Semantic Confidence Scoring and Rectification (SCSR) module further filters and corrects predictions. Experiments on Recipe1M show state-of-the-art performance and markedly improved semantic fidelity.

菜谱生成多模态语义准确

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。