解决视觉语言模型自进化中的简单任务偏倚问题
Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
- 通过重采样与分布重构,平衡简单与复杂任务的学习
- 在多个视觉推理任务上平均提升3.86分
- 适合希望提升模型复杂推理能力的研究者
自进化已成为提升大视觉语言模型(LVLMs)推理能力的主流范式,模型通过迭代探索和学习成功轨迹来优化自身。然而我们发现,在此过程中存在关键问题:模型擅长生成简单查询(头部数据)的高质量轨迹,却难以应对更复杂的任务(尾部数据)。这导致优化偏向简单推理技能,抑制了复杂推理能力的发展。随着迭代进行,这种不平衡日益加剧——我们称之为“马太效应”,最终造成性能瓶颈。为此,我们从分布重塑与轨迹重采样两个角度提出四种高效策略,实现探索-学习过程中的头尾再平衡。在Qwen2-VL-7B-Instruct和InternVL2.5-4B模型上的大量实验表明,所提方法在多个视觉推理任务中显著提升性能,相比原始自进化方法平均提升3.86分。
原文摘要 · Abstract (English)
Self-improvement has emerged as a mainstream paradigm for advancing the reasoning capabilities of large vision-language models (LVLMs), where models explore and learn from successful trajectories iteratively. However, we identify a critical issue during this process: the model excels at generating high-quality trajectories for simple queries (i.e., head data) but struggles with more complex ones (i.e., tail data). This leads to an imbalanced optimization that drives the model to prioritize simple reasoning skills, while hindering its ability to tackle more complex reasoning tasks. Over iterations, this imbalance becomes increasingly pronounced--a dynamic we term the "Matthew effect"--which ultimately hinders further model improvement and leads to performance bottlenecks. To counteract this challenge, we introduce four efficient strategies from two perspectives: distribution-reshaping and trajectory-resampling, to achieve head-tail re-balancing during the exploration-and-learning self-improvement process. Extensive experiments on Qwen2-VL-7B-Instruct and InternVL2.5-4B models across visual reasoning tasks demonstrate that our methods consistently improve visual reasoning capabilities, outperforming vanilla self-improvement by 3.86 points on average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。