arXiv:2411.01996cs.CLcs.AI2024-11

用新基准评估大模型在菜系融合中的创意与文化准确性

Culinary Class Wars: Evaluating LLMs using ASH in Cuisine Transfer Task

  • 设计ASH基准,从真实感、敏感度、协调性三维度评估菜系迁移
  • 发现大模型在文化细节处理上存在明显不足,生成结果常失真
  • 适合对跨文化美食生成感兴趣的AI研究者和食品科技开发者

大型语言模型(LLMs)在创意领域如烹饪艺术中展现出潜力,但面对将菜谱适配特定文化需求的任务时仍表现不佳。本研究聚焦于菜系迁移——将一种菜系元素融入另一种菜系——来评估LLMs的烹饪创造力。我们使用多种LLMs生成并评估文化适配的菜谱,并将其评价结果与人类及其它模型判断进行对比。为此,我们提出了ASH(真实性、敏感度、和谐性)基准,用于衡量LLMs在菜系迁移任务中的文化准确性和创造性。研究揭示了大模型在烹饪领域生成与评估能力上的关键洞见,凸显其在理解与应用文化细微差别方面的优势与局限。本项目所用代码与数据集将公开于 url{http://github.com/dmis-lab/CulinaryASH}。

原文摘要 · Abstract (English)

The advent of Large Language Models (LLMs) have shown promise in various creative domains, including culinary arts. However, many LLMs still struggle to deliver the desired level of culinary creativity, especially when tasked with adapting recipes to meet specific cultural requirements. This study focuses on cuisine transfer-applying elements of one cuisine to another-to assess LLMs' culinary creativity. We employ a diverse set of LLMs to generate and evaluate culturally adapted recipes, comparing their evaluations against LLM and human judgments. We introduce the ASH (authenticity, sensitivity, harmony) benchmark to evaluate LLMs' recipe generation abilities in the cuisine transfer task, assessing their cultural accuracy and creativity in the culinary domain. Our findings reveal crucial insights into both generative and evaluative capabilities of LLMs in the culinary domain, highlighting strengths and limitations in understanding and applying cultural nuances in recipe creation. The code and dataset used in this project will be openly available in \url{http://github.com/dmis-lab/CulinaryASH}.

菜系迁移大模型评估文化敏感性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。