动态更新的多模态评测基准,自动演化避免过时与数据污染。
MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

- 用多智能体自动化流水线持续生成新评测题。
- 每轮更新仅需30美元,1-2小时完成,新增5.9K题。
- 保持模型排名稳定,减少记忆偏差,适合长期评测研究。
评估基准对视觉语言模型至关重要,但多数多模态基准为静态,易受时间过时、数据污染和维护成本影响。我们提出 MMBench-Live,一个由多智能体驱动的自动化流水线构建的持续演化的多模态基准。该框架将基准演化视为任务引导的数据集构建,整合结构化基准定义、反馈控制的实时数据采集以及可验证的问答生成与可执行推理。为保持跨版本可比性,引入分布一致的更新策略,从原始基准提取任务相关视觉模式以指导数据收集与筛选。基于 MMBench 构建,MMBench-Live 包含 5.9K 个新生成的评估实例,答案正确率高,每次更新成本约 30 美元,耗时 1-2 小时。大量实验表明,其保持了稳定的模型排名,与原基准语义对齐良好,且表现出更弱的数据污染记忆信号,展示出可持续多模态基准演化的实用与可扩展范式。项目地址:https://github.com/PRIS-CV/MMBench-Live。
原文摘要 · Abstract (English)
Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance. We present MMBench-Live, a continuously evolving multimodal benchmark built by a multi-agent-driven automated pipeline. Our framework treats benchmark evolution as task-guided dataset construction, integrating structured benchmark specification, feedback-controlled real-time data acquisition, and verifiable QA generation with executable reasoning. To maintain cross-version comparability, we introduce a distribution-consistent update strategy that extracts task-related visual patterns from the original benchmark to guide data collection and filtering. Instantiated from MMBench, MMBench-Live contains 5.9K newly generated evaluation instances with a high answer correctness rate, while each update costs about USD 30 and takes 1-2 hours. Extensive evaluations show that MMBench-Live preserves stable model rankings, maintains semantic alignment with the original benchmark, and exhibits weaker contamination-related memorization signals, suggesting a practical and scalable paradigm for sustainable multimodal benchmark evolution. The project is available at https://github.com/PRIS-CV/MMBench-Live.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。