arXiv:2602.22601cs.LGcs.CV2026-02中稿 · CVPR

提出新方法缓解多模态大模型持续学习中的偏见问题。

$ϕ$-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal Models

  • 基于偏好优化构建新范式,缓解灾难性遗忘。
  • 在多个基准上超越现有方法,提升公平性与性能。
  • 适合关注模型公平性与持续学习的研究者。

大型多模态模型(LMMs)在持续学习中的公平性是一个新兴但研究不足的挑战,尤其在数据分布不均衡时可能导致模型更新偏差和任务间性能下降。尽管近期研究在缓解灾难性遗忘方面取得进展,但由数据不平衡引发的公平性问题仍缺乏深入探索。本文提出一种新的公平性直接偏好优化(FaiDPO,即ϕ-DPO)框架用于LMMs的持续学习。首先,基于直接偏好优化(DPO)提出一种新范式,通过配对偏好信号对齐学习过程,减轻遗忘问题;其次,识别传统DPO在数据不均衡下的局限,设计ϕ-DPO损失函数以显式纠正分布偏差。理论分析证明该方法同时缓解遗忘与数据不平衡。此外,为支持ϕ-DPO,我们在现有基准上构建了持续学习场景下的配对偏好标注。大量实验与消融研究显示,所提ϕ-DPO在多个基准上达到最先进性能,显著优于现有LMM持续学习方法。

原文摘要 · Abstract (English)

Fairness in Continual Learning for Large Multimodal Models (LMMs) is an emerging yet underexplored challenge, particularly in the presence of imbalanced data distributions that can lead to biased model updates and suboptimal performance across tasks. While recent continual learning studies have made progress in addressing catastrophic forgetting, the problem of fairness caused the imbalanced data remains largely underexplored. This paper presents a novel Fairness Direct Preference Optimization (FaiDPO or $ϕ$-DPO) framework for continual learning in LMMs. In particular, we first propose a new continual learning paradigm based on Direct Preference Optimization (DPO) to mitigate catastrophic forgetting by aligning learning with pairwise preference signals. Then, we identify the limitations of conventional DPO in imbalanced data and present a new $ϕ$-DPO loss that explicitly addresses distributional biases. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both forgetting and data imbalance. Additionally, to enable $ϕ$-DPO-based continual learning, we construct pairwise preference annotations for existing benchmarks in the context of continual learning. Extensive experiments and ablation studies show the proposed $ϕ$-DPO achieves State-of-the-Art performance across multiple benchmarks, outperforming prior continual learning methods of LMMs.

持续学习多模态公平性偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。