arXiv:2602.01640cs.CL2026-02被引 3

用智能代理自动构建评估体系,让机器人视觉模型测试更高效准确

A2Eval: Agentic and Automated Evaluation for Embodied Brain

  • 设计两个协作智能体,自动发现能力维度并生成平衡评估集
  • 评估集压缩85%,计算成本降77%,速度提升4.6倍,质量不下降
  • 纠正排名偏差,与人类评价相关性达0.85,适合研究者快速迭代模型

当前具身视觉语言模型评估依赖静态、人工标注的基准,存在严重冗余和覆盖不均问题。该范式耗时耗力,浪费计算与标注资源,推高成本并扭曲模型排名,阻碍迭代开发。为此,我们提出首个智能自动化评估框架A2Eval,通过两个协同智能体实现基准自动生成与评估。数据代理自主识别能力维度并构建均衡紧凑的评估套件,评估代理合成并验证可执行评估流程,实现完全自动化、高保真度评估。在10个基准和13个模型上验证,A2Eval将评估集压缩85%,整体计算成本降低77%,速度提升4.6倍,同时保持评估质量。关键的是,它纠正了系统性排名偏差,人类对齐性达斯皮尔曼相关系数0.85,排名保真度为肯德尔等级相关系数0.81,确立了高保真、低成本具身评估新标准。代码与数据将很快公开。

原文摘要 · Abstract (English)

Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm drains computational and annotation resources, inflates costs, and distorts model rankings, ultimately stifling iterative development. To address this, we propose Agentic Automatic Evaluation (A2Eval), the first agentic framework that automates benchmark curation and evaluation through two collaborative agents. The Data Agent autonomously induces capability dimensions and assembles a balanced, compact evaluation suite, while the Eval Agent synthesizes and validates executable evaluation pipelines, enabling fully autonomous, high-fidelity assessment. Evaluated across 10 benchmarks and 13 models, A2Eval compresses evaluation suites by 85%, reduces overall computational costs by 77%, and delivers a 4.6x speedup while preserving evaluation quality. Crucially, A2Eval corrects systematic ranking biases, improves human alignment to Spearman's rho=0.85, and maintains high ranking fidelity (Kendall's tau=0.81), establishing a new standard for high-fidelity, low-cost embodied assessment. Our code and data will be public soon.

具身智能自动评估智能体视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。