提升视觉语言模型长视频与高分辨率图像理解能力
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models

- 采用自动降级采样与图像区域保留技术,保持上下文完整性和视觉细节
- 在512帧视频上达到72.4%的Video-MME得分,媲美GPT-4o等顶级模型
- 适用于长视频理解、高分辨率图像分析场景,适合多模态研究者使用
我们提出Eagle 2.5,一套面向长上下文多模态学习的前沿视觉语言模型(VLMs)家族。针对长视频理解与高分辨率图像解析的挑战,引入通用框架统一处理两类任务。训练框架包含自动降级采样(Automatic Degrade Sampling)和图像区域保留(Image Area Preservation)技术,有效维护上下文连贯性与视觉细节。同时,优化长上下文数据训练流程以提升效率。此外,我们构建了Eagle-Video-110K数据集,融合故事级与片段级标注,支持长视频理解。Eagle 2.5在长上下文多模态基准测试中表现显著提升,其最佳模型Eagle 2.5-8B在512帧输入下于Video-MME上达到72.4%,与GPT-4o及Qwen2.5-VL-72B、InternVL2.5-78B等大模型性能相当。
原文摘要 · Abstract (English)
We introduce Eagle 2.5, a family of frontier vision-language models (VLMs) for long-context multimodal learning. Our work addresses the challenges in long video comprehension and high-resolution image understanding, introducing a generalist framework for both tasks. The proposed training framework incorporates Automatic Degrade Sampling and Image Area Preservation, two techniques that preserve contextual integrity and visual details. The framework also includes numerous efficiency optimizations in the pipeline for long-context data training. Finally, we propose Eagle-Video-110K, a novel dataset that integrates both story-level and clip-level annotations, facilitating long-video understanding. Eagle 2.5 demonstrates substantial improvements on long-context multimodal benchmarks, providing a robust solution to the limitations of existing VLMs. Notably, our best model Eagle 2.5-8B achieves 72.4% on Video-MME with 512 input frames, matching the results of top-tier commercial model such as GPT-4o and large-scale open-source models like Qwen2.5-VL-72B and InternVL2.5-78B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。