arXiv:2608.29249cs.AIcs.CL2026-08

为大模型生成的印度菜谱提供可审计的验证框架,防幻觉、保准确。

Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge

论文配图:Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge
图 1 · 摘自论文原文
  • 用语法、词汇、统计和检索多阶段检测菜谱错误
  • 在12类常见问题中识别出87%的结构与逻辑错误
  • 适合构建食品知识图谱或开发多语言菜谱系统的人看

线上餐饮生态中,越来越多的食谱由大语言模型(LLMs)生成、修改或总结。尽管内容看似合理,但常包含虚构食材、错误分量或文化不合理的搭配,限制了其在下游应用和知识图谱构建中的适用性。本文提出一种半自动化的声音度评估流程,用于验证从非正式来源提取并经大模型增强的结构化食谱数据。该流程作为FKG.in(印度食物知识图谱)的一部分,通过多阶段方法识别并修复常见缺陷,包括结构不一致、语义与逻辑矛盾,以及与原文偏差,结合形式语法、词汇检查、统计启发式、基于Set Transformer的连贯性建模和检索验证。尽管在印度菜谱上进行评估,所提方法可扩展至更广泛的多语言、跨文化烹饪领域。本文提供了一个实用、可审计且与应用场景无关的验证框架,强化了大模型时代可机器读取的食品知识基础设施。

原文摘要 · Abstract (English)

The online culinary ecosystem is increasingly populated by recipe content generated, modified, or summarized by Large Language Models (LLMs). While often plausible, such outputs may contain hallucinated ingredients, misrepresented quantities, or culturally implausible combinations, limiting their suitability for downstream applications and knowledge graph construction. In this paper, we present a semi-automated soundness assessment workflow for validating structured recipe data extracted and augmented by LLMs from informal culinary sources. Developed as part of FKG(.in), a knowledge graph of Indian food, the pipeline identifies and addresses common failure modes, including structural inconsistencies, semantic and logical incoherence, and deviations from the source text, through a multi-stage process combining formal grammars, vocabulary-based checks, statistical heuristics, Set Transformer-based coherence modeling, and retrieval-based verification. Although evaluated on Indian recipes, the proposed methods are applicable to broader multilingual and multicultural culinary domains. We provide a practical, auditable, and application-agnostic framework for validating LLM-augmented recipe data, thereby strengthening the foundations of machine-readable food knowledge infrastructures in the era of LLM-generated content.

知识图谱大模型验证菜谱生成印度饮食

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。