用多样化提示+模型协作,让大模型在视频理解上表现超越此前最佳方法。
Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles
- 设计多种提示风格与处理范式,引导大模型注意力。
- 单一模型直接使用已超越之前最优方法,性能提升显著。
- 引入周期结果协作阶段,实现模型间互补,效果更优。
本文介绍在CVPR 2025年Ego4D EgoSchema挑战赛中的入围方案(确认于2025年5月20日)。受大型模型成功启发,我们评估并利用当前可获取的多模态大模型,通过少样本学习与模型集成策略适配视频理解任务。具体地,系统探索并评估了多样化的提示风格与处理范式,有效引导大模型注意力,充分释放其强大的泛化与适应能力。实验表明,采用精心设计的方法后,仅使用单一多模态模型即已超越此前的最先进(SOTA)方法,该方法包含多个额外处理步骤。此外,进一步引入一个促进周期性结果协作与集成的阶段,带来显著性能提升。我们希望本工作能为大模型的实际应用提供参考,并激发未来研究。代码已开源:https://github.com/XiongjunGuan/EgoSchema-CVPR25。
原文摘要 · Abstract (English)
In this paper, we present the runner-up solution for the Ego4D EgoSchema Challenge at CVPR 2025 (Confirmed on May 20, 2025). Inspired by the success of large models, we evaluate and leverage leading accessible multimodal large models and adapt them to video understanding tasks via few-shot learning and model ensemble strategies. Specifically, diversified prompt styles and process paradigms are systematically explored and evaluated to effectively guide the attention of large models, fully unleashing their powerful generalization and adaptability abilities. Experimental results demonstrate that, with our carefully designed approach, directly utilizing an individual multimodal model already outperforms the previous state-of-the-art (SOTA) method which includes several additional processes. Besides, an additional stage is further introduced that facilitates the cooperation and ensemble of periodic results, which achieves impressive performance improvements. We hope this work serves as a valuable reference for the practical application of large models and inspires future research in the field. Our Code is available at https://github.com/XiongjunGuan/EgoSchema-CVPR25.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。