arXiv:2510.20393cs.CVcs.MM2025-10被引 1

解决跨模态食物图文检索中的文化偏见问题

Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval

  • 基于因果推理预测图像中缺失的烹饪细节并注入表示学习
  • 在多语言多文化数据集上提升细粒度食材与做法检索效果
  • 适合关注跨文化美食生成与精准检索的研究者

现有图像到菜谱检索方法隐含假设:一张食物图像能完整体现菜谱中的文字信息。但图像仅反映成品视觉外观,无法呈现烹饪过程。因此,跨模态表示学习容易偏向捕捉主导视觉特征,忽略不显性的、对检索至关重要的细微差异(如配料或烹饪方式)。当训练数据包含多种菜系时,这种偏差更为严重。本文提出一种新的因果方法:预测图像中可能被忽略的烹饪元素,并将其显式注入跨模态表示学习以缓解偏差。在标准单语Recipe1M数据集和新构建的多语言多文化菜系数据集上进行实验,结果表明该方法能有效识别细微配料与烹饪动作,在单语和多语多文化数据集上均取得优异检索性能。

原文摘要 · Abstract (English)

Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and not the underlying cooking process. Consequently, learning cross-modal representations to bridge the modality gap between images and recipes tends to ignore subtle, recipe-specific details that are not visually apparent but are crucial for recipe retrieval. Specifically, the representations are biased to capture the dominant visual elements, resulting in difficulty in ranking similar recipes with subtle differences in use of ingredients and cooking methods. The bias in representation learning is expected to be more severe when the training data is mixed of images and recipes sourced from different cuisines. This paper proposes a novel causal approach that predicts the culinary elements potentially overlooked in images, while explicitly injecting these elements into cross-modal representation learning to mitigate biases. Experiments are conducted on the standard monolingual Recipe1M dataset and a newly curated multilingual multicultural cuisine dataset. The results indicate that the proposed causal representation learning is capable of uncovering subtle ingredients and cooking actions and achieves impressive retrieval performance on both monolingual and multilingual multicultural datasets.

跨模态检索食物生成因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。