arXiv:2511.15201cs.CVcs.MM2025-11

用因果模型消除食物图文检索中的偏见,提升精准度。

Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval

  • 引入因果理论识别配料为混淆因子,通过反向门调整去偏。
  • 在Recipe1M数据集上实现MedR=1的最优检索性能,覆盖多种测试规模。
  • 提出可即插即用的多标签配料分类模块,适合图像-食谱跨模态研究者。

本文针对食谱与食物图像跨模态检索中的表示学习挑战展开研究。由于食谱与其成品之间存在因果关系,现有方法将食谱视为描述菜品视觉外观的文本源,会引入偏差,导致图像与食谱相似性判断失准。因烹饪过程、摆盘方式及拍摄条件等因素,图像未必完整呈现食谱细节,当前表示学习倾向于捕捉主导的视觉-文本对齐,忽略影响检索相关性的细微差异。本文基于因果理论建模该偏差,指出配料是主要混淆因子,简单反向门调整即可缓解偏差。通过因果干预,重构传统食物到食谱检索模型,增加去偏项。基于此理论指导的框架,在Recipe1M数据集上验证了不同测试规模(1K、10K、50K)下均达到MedR=1的基准性能。同时提出一个即插即用的神经模块——多标签配料分类器,用于去偏。在Recipe1M上实现了新的最佳检索效果。

原文摘要 · Abstract (English)

This paper addresses the challenges of learning representations for recipes and food images in the cross-modal retrieval problem. As the relationship between a recipe and its cooked dish is cause-and-effect, treating a recipe as a text source describing the visual appearance of a dish for learning representation, as the existing approaches, will create bias misleading image-and-recipe similarity judgment. Specifically, a food image may not equally capture every detail in a recipe, due to factors such as the cooking process, dish presentation, and image-capturing conditions. The current representation learning tends to capture dominant visual-text alignment while overlooking subtle variations that determine retrieval relevance. In this paper, we model such bias in cross-modal representation learning using causal theory. The causal view of this problem suggests ingredients as one of the confounder sources and a simple backdoor adjustment can alleviate the bias. By causal intervention, we reformulate the conventional model for food-to-recipe retrieval with an additional term to remove the potential bias in similarity judgment. Based on this theory-informed formulation, we empirically prove the oracle performance of retrieval on the Recipe1M dataset to be MedR=1 across the testing data sizes of 1K, 10K, and even 50K. We also propose a plug-and-play neural module, which is essentially a multi-label ingredient classifier for debiasing. New state-of-the-art search performances are reported on the Recipe1M dataset.

跨模态因果推理去偏食谱检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。