提出RFHNet网络,提升细粒度食物图像检索精度。
RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image Retrieval

- 分层结构捕捉全局与局部视觉细节
- 12位哈希下平均精度提升4.44%至17.20%
- 适合智能餐饮与饮食监测场景
细粒度食物图像检索是计算美食学的关键任务,应用于食品溯源、膳食监控和智慧餐饮系统。尽管基于哈希的检索因存储效率高和汉明距离计算快而适用于大规模搜索,但现有方法在细粒度食物场景中表现不佳,因细微局部语义和频率敏感视觉线索至关重要。为此,我们提出RFHNet,一种级联分层哈希网络,通过多层级表征同时捕捉全局结构与细粒度局部细节。RFHNet包含三个组件:(1) 细粒度关系建模(FRM)以捕获相似食物成分间的细微视觉差异;(2) 多频调制融合(MFMF)提取信息丰富的多频特征;(3) 分层语义协同(HSS)自适应整合多层级表示并生成判别性哈希码。在六个食物专用基准上的实验表明,RFHNet始终优于当前最优哈希方法,在12位哈希下平均精度提升4.44%至17.20%。结果验证了RFHNet在大规模视觉食物检索及智慧餐饮应用中的有效性。源代码将在发表后公开。
原文摘要 · Abstract (English)
Fine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart catering systems. Although hashing-based retrieval is attractive for large-scale search due to its storage efficiency and fast Hamming-distance computation, existing methods often perform poorly in fine-grained food scenarios, where subtle local semantics and frequency-sensitive visual cues are essential. To address this challenge, we propose RFHNet, a cascaded hierarchical hashing network that captures both global structure and fine-grained local details through multi-level representations. RFHNet includes three components: (1) Fine-grained Relation Modeling (FRM) to capture subtle visual differences among similar food components; (2) Multi-Frequency Modulated Fusion (MFMF) to extract informative multi-frequency features; and (3) Hierarchical Semantic Synergy (HSS) to adaptively integrate multi-level representations and generate discriminative hash codes. Experiments on six food-specific benchmarks show that RFHNet consistently outperforms state-of-the-art hashing methods, with mAP gains of 4.44\% to 17.20\% at 12 bits. These results validate the effectiveness of RFHNet for large-scale visual food retrieval and smart catering applications. The source code will be released upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。