用拓扑输运损失生成更准的菜谱,食材和步骤都更靠谱。
Losses that Cook: Topological Optimal Transport for Structured Recipe Generation
- 将食材列表视为嵌入空间中的点云,用拓扑损失优化预测与真实食材的差异。
- 在食材、动作、时间、温度等指标上均有显著提升,尤其时温精度提高37%。
- 适合需要高精度流程和成分匹配的生成任务,如智能食谱系统。
菜谱是复杂的程序性文本,不仅要求语言流畅准确,还需时间、温度、步骤连贯性以及正确配料组合。现有训练方法多依赖交叉熵,仅关注语言流畅性。基于RECIPE-NLG,我们研究多种复合目标函数,提出一种新的拓扑损失,将食材列表表示为嵌入空间中的点云,最小化预测与真实食材之间的分布差异。通过标准NLG指标与菜谱特有指标评估,发现该损失显著提升食材与动作层面的表现。Dice损失在时间/温度精度上表现优异,混合损失在量与时间间取得良好平衡,并实现协同增益。人类偏好分析显示,模型在62%情况下更受青睐。
原文摘要 · Abstract (English)
Cooking recipes are complex procedures that require not only a fluent and factual text, but also accurate timing, temperature, and procedural coherence, as well as the correct composition of ingredients. Standard training procedures are primarily based on cross-entropy and focus solely on fluency. Building on RECIPE-NLG, we investigate the use of several composite objectives and present a new topological loss that represents ingredient lists as point clouds in embedding space, minimizing the divergence between predicted and gold ingredients. Using both standard NLG metrics and recipe-specific metrics, we find that our loss significantly improves ingredient- and action-level metrics. Meanwhile, the Dice loss excels in time/temperature precision, and the mixed loss yields competitive trade-offs with synergistic gains in quantity and time. A human preference analysis supports our finding, showing our model is preferred in 62% of the cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。