用多模态知识图谱提升食物问答的准确与多样性。
Enhancing Food-Domain Question Answering with a Multimodal Knowledge Graph: Hybrid QA Generation and Diversity Analysis
- 构建包含1.4万张图的食品多模态知识图谱,融合生成式AI。
- 联合微调后模型在多项指标上提升超30%,图像生成更真实。
- 适合做食品智能问答、个性化食谱推荐的研究与开发者。
我们提出一个统一的食物领域问答框架,结合大规模多模态知识图谱(MMKG)与生成式AI。该MMKG关联了13,000个食谱、3,000种食材、140,000条关系和14,000张图像。基于40个模板并使用LLaVA/DeepSeek增强,生成40,000组问答对。对Meta LLaMA 3.1-8B与Stable Diffusion 3.5-Large进行联合微调后,BERTScore提升16.2%,FID降低37.8%,CLIP对齐度提高31.1%。诊断分析包括基于CLIP的错配检测(从35.2%降至7.3%)和基于LLaVA的幻觉检查,保障事实性与视觉真实性。混合检索-生成策略实现94.1%的图像复用准确率和85%的合成质量适配度。结果表明,结构化知识与多模态生成协同可显著提升食物问答的可靠性与多样性。
原文摘要 · Abstract (English)
We propose a unified food-domain QA framework that combines a large-scale multimodal knowledge graph (MMKG) with generative AI. Our MMKG links 13,000 recipes, 3,000 ingredients, 140,000 relations, and 14,000 images. We generate 40,000 QA pairs using 40 templates and LLaVA/DeepSeek augmentation. Joint fine-tuning of Meta LLaMA 3.1-8B and Stable Diffusion 3.5-Large improves BERTScore by 16.2\%, reduces FID by 37.8\%, and boosts CLIP alignment by 31.1\%. Diagnostic analyses-CLIP-based mismatch detection (35.2\% to 7.3\%) and LLaVA-driven hallucination checks-ensure factual and visual fidelity. A hybrid retrieval-generation strategy achieves 94.1\% accurate image reuse and 85\% adequacy in synthesis. Our results demonstrate that structured knowledge and multimodal generation together enhance reliability and diversity in food QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。