解决食物多模态学习中模型遗忘问题,提升持续学习能力。
Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning
- 用双低秩适配器分离任务特有与共享知识。
- 在Uni-Food数据集上显著降低遗忘率,效果优于现有方法。
- 适合做健康营养、慢性病预防的持续学习研究者。
食物分析在个性化营养和慢性病预防等健康任务中日益重要。然而,现有的大型多模态模型(LMMs)在学习新任务时面临灾难性遗忘,需从头重新训练,成本高昂。为此,我们提出一种新型多模态食物持续学习框架,结合双LoRA架构与质量增强的伪回放策略。针对每个任务引入两个互补的低秩适配器:专用LoRA通过正交约束学习任务特异性知识,避免干扰旧任务;协作LoRA通过伪回放整合跨任务共享知识。为提高回放数据可靠性,提出质量增强的伪回放策略,利用自一致性与语义相似性减少生成样本中的幻觉。在综合性Uni-Food数据集上的实验表明,该方法有效缓解遗忘,是首个适用于复杂食物任务的有效持续学习方案。
原文摘要 · Abstract (English)
Food analysis has become increasingly critical for health-related tasks such as personalized nutrition and chronic disease prevention. However, existing large multimodal models (LMMs) in food analysis suffer from catastrophic forgetting when learning new tasks, requiring costly retraining from scratch. To address this, we propose a novel continual learning framework for multimodal food learning, integrating a Dual-LoRA architecture with Quality-Enhanced Pseudo Replay. We introduce two complementary low-rank adapters for each task: a specialized LoRA that learns task-specific knowledge with orthogonal constraints to previous tasks' subspaces, and a cooperative LoRA that consolidates shared knowledge across tasks via pseudo replay. To improve the reliability of replay data, our Quality-Enhanced Pseudo Replay strategy leverages self-consistency and semantic similarity to reduce hallucinations in generated samples. Experiments on the comprehensive Uni-Food dataset show superior performance in mitigating forgetting, representing the first effective continual learning approach for complex food tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。