首个马其顿菜谱数据集,助力小语种饮食文化研究
Building a Macedonian Recipe Dataset: Collection, Parsing, and Comparative Analysis
- 通过网络爬取与结构化解析构建马其顿菜谱数据集
- 发现马其顿菜系中独特的食材组合模式
- 为小语种饮食文化研究提供新资源,适合语言学与食品科学交叉研究者
计算烹饪学越来越依赖多样且高质量的菜谱数据集来捕捉区域饮食传统。尽管主流语言已有大规模数据集,马其顿菜谱在数字研究中仍严重不足。本文首次系统性地通过网络爬取和结构化解析构建马其顿菜谱数据集,解决了成分描述异构性问题,包括单位、数量和修饰词的标准化。基于点互信息(Pointwise Mutual Information)和提升度(Lift score)等指标的探索性分析,揭示了体现马其顿菜肴特色的食材共现模式。该数据集为研究未充分代表语言的饮食文化提供了新资源,并揭示了马其顿烹饪传统的独特规律。
原文摘要 · Abstract (English)
Computational gastronomy increasingly relies on diverse, high-quality recipe datasets to capture regional culinary traditions. Although there are large-scale collections for major languages, Macedonian recipes remain under-represented in digital research. In this work, we present the first systematic effort to construct a Macedonian recipe dataset through web scraping and structured parsing. We address challenges in processing heterogeneous ingredient descriptions, including unit, quantity, and descriptor normalization. An exploratory analysis of ingredient frequency and co-occurrence patterns, using measures such as Pointwise Mutual Information and Lift score, highlights distinctive ingredient combinations that characterize Macedonian cuisine. The resulting dataset contributes a new resource for studying food culture in underrepresented languages and offers insights into the unique patterns of Macedonian culinary tradition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。