arXiv:2506.09081cs.CVcs.AI2025-06ACL被引 5

一个开源多模态模型评估框架,支持多种视觉语言任务的高效测试。

FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation

论文配图:FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
图 1 · 摘自论文原文
  • 分离推理与评估,支持灵活资源调度和任务扩展
  • 采用vLLM等加速工具,评估效率显著提升
  • 适合研究者快速评测多模态模型性能

我们提出FlagEvalMM,一个开源评估框架,用于全面评估多模态模型在视觉语言理解与生成任务中的表现,涵盖视觉问答、文本到图像/视频生成、图文检索等。该框架通过独立的评估服务将模型推理与评估解耦,实现灵活资源配置,并支持新任务与新模型的无缝集成。同时,利用vLLM、SGLang等先进推理加速工具及异步数据加载技术,大幅提升评估效率。大量实验表明,FlagEvalMM能准确高效地揭示模型的优势与局限,是推动多模态研究的重要工具。框架已公开,地址为https://github.com/flageval-baai/FlagEvalMM。

原文摘要 · Abstract (English)

We present FlagEvalMM, an open-source evaluation framework designed to comprehensively assess multimodal models across a diverse range of vision-language understanding and generation tasks, such as visual question answering, text-to-image/video generation, and image-text retrieval. We decouple model inference from evaluation through an independent evaluation service, thus enabling flexible resource allocation and seamless integration of new tasks and models. Moreover, FlagEvalMM utilizes advanced inference acceleration tools (e.g., vLLM, SGLang) and asynchronous data loading to significantly enhance evaluation efficiency. Extensive experiments show that FlagEvalMM offers accurate and efficient insights into model strengths and limitations, making it a valuable tool for advancing multimodal research. The framework is publicly accessible at https://github.com/flageval-baai/FlagEvalMM.

多模态评估开源框架推理加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。