用课程引导自进化,让视频理解模型逐步提升能力
CurEvo: Curriculum-Guided Self-Evolution for Video Understanding

- 根据模型水平动态调整任务难度,形成有节奏的学习路径
- 在4个视频问答基准上,准确率和语义评分均显著提升
- 适合希望实现自主进化的视频理解研究者
近期自进化视频理解框架展现了无需人工标注的自主学习潜力。然而,现有方法普遍存在优化控制弱、难度进展无序的问题,缺乏迭代学习过程中的结构化引导。为此,我们提出CurEvo,一种引入课程学习的自进化框架,通过动态调节任务难度、优化评估标准并平衡数据多样性,依据模型能力构建课程引导的反馈循环,使学习复杂度与模型能力对齐。基于此,我们设计多维度自适应问答框架,协同进化感知、识别与理解维度的提问生成与答案评估,确保连贯可测的课程演进。该整合将弱控制的自进化转变为更结构化的学习过程。在7个骨干网络上,CurEvo在4个VideoQA基准上均持续提升基准准确率与评估器语义得分,验证了课程引导自进化在视频理解中的有效性。
原文摘要 · Abstract (English)
Recent advances in self-evolution video understanding frameworks have demonstrated the potential of autonomous learning without human annotations. However, existing methods often suffer from weakly controlled optimization and uncontrolled difficulty progression, as they lack structured guidance throughout the iterative learning process. To address these limitations, we propose CurEvo, a curriculum-guided self-evolution framework that introduces curriculum learning into self-evolution to achieve more structured and progressive model improvement. CurEvo dynamically regulates task difficulty, refines evaluation criteria, and balances data diversity according to model competence, forming a curriculum-guided feedback loop that aligns learning complexity with model capability. Built upon this principle, we develop a multi-dimensional adaptive QA framework that jointly evolves question generation and answer evaluation across perception, recognition, and understanding dimensions, ensuring coherent and measurable curriculum progression. Through this integration, CurEvo transforms weakly controlled self-evolution into a more structured learning process for autonomous video understanding. Across seven backbones, CurEvo consistently improves both benchmark accuracy and evaluator-based semantic score on four VideoQA benchmarks, validating the effectiveness of curriculum-guided self-evolution for video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。