新基准MiroEval全面评估多模态研究智能体的全过程与结果表现。
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
- 构建100个真实任务,支持文本与多模态混合,可定期更新
- 三维度评估:结果质量、事实真实性、研究过程,发现过程越优结果越好
- 多模态任务难度显著提升,多数系统性能下降3-10分
深度研究系统进展迅速,但评估仍滞后于实际需求。现有基准多仅基于固定评分标准评估最终报告,忽视研究过程;且多模态覆盖有限,依赖合成任务,无法随知识演进更新。为此,我们提出MiroEval,一个面向深度研究系统的基准与评估框架。该基准包含100项任务(70项文本类,30项多模态),均源于真实用户需求,通过双路径流程构建,支持周期性更新,实现动态演进。评估体系涵盖三个互补维度:针对任务特性的自适应合成质量评估、基于主动检索与推理的事实性验证(覆盖网页与多模态附件)、以及以过程为中心的审计(考察搜索、推理与迭代优化)。对13个系统的评估揭示三大发现:三个维度捕获互补能力特征,各显系统优劣;过程质量能可靠预测整体表现,揭示输出指标未捕捉的缺陷;多模态任务挑战更大,多数系统性能下降3至10分。MiroThinker系列表现最均衡,其中MiroThinker-H1在两类设置下均排名第一。人工验证与鲁棒性测试确认了基准可靠性。MiroEval为下一代深度研究智能体提供全方位诊断工具。
原文摘要 · Abstract (English)
Recent progress in deep research systems has been impressive, but evaluation still lags behind real user needs. Existing benchmarks predominantly assess final reports using fixed rubrics, failing to evaluate the underlying research process. Most also offer limited multimodal coverage, rely on synthetic tasks that do not reflect real-world query complexity, and cannot be refreshed as knowledge evolves. To address these gaps, we introduce MiroEval, a benchmark and evaluation framework for deep research systems. The benchmark comprises 100 tasks (70 text-only, 30 multimodal), all grounded in real user needs and constructed via a dual-path pipeline that supports periodic updates, enabling a live and evolving setting. The proposed evaluation suite assesses deep research systems along three complementary dimensions: adaptive synthesis quality evaluation with task-specific rubrics, agentic factuality verification via active retrieval and reasoning over both web sources and multimodal attachments, and process-centric evaluation audits how the system searches, reasons, and refines throughout its investigation. Evaluation across 13 systems yields three principal findings: the three evaluation dimensions capture complementary aspects of system capability, with each revealing distinct strengths and weaknesses across systems; process quality serves as a reliable predictor of overall outcome while revealing weaknesses invisible to output-level metrics; and multimodal tasks pose substantially greater challenges, with most systems declining by 3 to 10 points. The MiroThinker series achieves the most balanced performance, with MiroThinker-H1 ranking the highest overall in both settings. Human verification and robustness results confirm the reliability of the benchmark and evaluation framework. MiroEval provides a holistic diagnostic tool for the next generation of deep research agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。