用推理链提升印度菜视觉问答准确率,平均提升10个百分点。
Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
- 构建自动验证的推理链,引导模型分步理解复杂厨艺背景。
- 在印度菜视觉问答上,加入推理链后准确率平均提升10个百分点。
- 适合关注多步推理与小模型微调的研究者和应用开发者。
印度菜系文化多样,现有视觉问答系统多偏重西方食物,存在明显局限。尽管已有针对印度菜的VQA数据集出现,但其方法仍为两步流程:先生成答案,再解释理由。本文认为,印度菜VQA需多步推理以理解复杂的烹饪语境及食物间关系。为此,我们基于问答对自动生成推理链,仅需少量人工干预。通过使用自动验证的推理链微调小型大语言模型和视觉语言模型,并借助强化学习在更大规模数据上进一步训练,显著提升了性能。实验显示,引入推理链后,基线模型平均准确率提升10个百分点。本文还深入分析了推理链对印度菜视觉问答任务的影响。关键词:FoodVQA,推理链,强化学习,知识图谱。
原文摘要 · Abstract (English)
The immense diversity in the culture and culinary of Indian cuisines calls attention to the major shortcoming of the existing Visual Question Answering(VQA) systems which are inclined towards the foods from Western region. Recent attempt towards building a VQA dataset for Indian food is a step towards addressing this challenge. However, their approach towards VQA follows a two-step process in which the answer is generated first, followed by the explanation of the expected answer. In this work, we claim that food VQA requires to follow a multi-step reasoning process to arrive at an accurate answer, especially in the context of India food, which involves understanding complex culinary context and identifying relationships between various food items. With this hypothesis we create reasoning chains upon the QA with minimal human intervention. We fine-tune smaller LLMs and VLMs with auto-validated reasoning chains and further train them using reinforcement learning with larger data. With augmentation of reasoning chains, we observed accuracy improvement of an average 10 percentage points on the baseline. We provide detailed analysis in terms the effect of addition of reasoning chains for the Indian Food VQA task. Index Terms - FoodVQA, Reasoning Chains, Reinforcement Learning, Knowledge Graph.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。