arXiv:2605.19986cs.ROcs.CV2026-05

提出诊断框架MetaFine,揭示精细操作中被传统评估掩盖的短板。

Beyond Binary Success: A Diagnostic Meta-Evaluation Framework for Fine-Grained Manipulation

论文配图:Beyond Binary Success: A Diagnostic Meta-Evaluation Framework for Fine-Grained Manipulation
图 1 · 摘自论文原文
  • 构建三轴诊断框架,拆解理解、感知、行为控制能力
  • 发现视觉编码器保持局部空间结构是精度瓶颈,改进后解锁新能力
  • 支持真实与仿真混合验证,助力模型可复现优化

精细操作要求局部属性定位、高保真空间感知与约束遵守的运动执行紧密耦合,但现有具身AI基准将这些能力压缩为二元成功率,导致性能报告虚高达70%,并掩盖了阻碍实际部署的架构瓶颈。我们提出MetaFine诊断性元评估框架,从理解、感知和可控行为三个维度解耦操作能力。基于组合任务图,该框架整合异构外部基准,统一协议下重构复杂度可调的诊断场景。对前沿视觉-语言-动作(VLA)模型的评估揭示了传统指标无法捕捉的维度特异性失败。通过定向因果干预,我们确认视觉编码器保持局部空间结构的能力是精细精度的关键瓶颈:提升此能力可直接释放此前无法实现的操作能力,无需修改下游策略。MetaFine还支持混合真实-仿真验证,利用有限真实回放校准可扩展的仿真估计,实现更稳定的物理基准测试。通过将评估从排名转向诊断,MetaFine使基准测试成为修复物理灵巧性底层能力的行动指南。框架、基准与资源将公开发布于项目页:https://metafine.github.io/。

原文摘要 · Abstract (English)

Fine-grained manipulation marks a regime where global scene context no longer suffices, and success hinges on the tight coupling of local attribute grounding, high-fidelity spatial perception, and constraint-respecting motor execution. However, current embodied AI benchmarks collapse these capacities into binary success rates, systematically inflating reported capabilities by up to 70% and masking the architectural bottlenecks that impede real-world deployment. We introduce MetaFine, a diagnostic meta-evaluation framework that disentangles manipulation competency along three axes: understanding, perception, and controlled behavior. Built on a compositional task graph, MetaFine absorbs heterogeneous external benchmarks and reconstructs them into diagnostic scenarios of varying complexity under a unified protocol. Evaluating state-of-the-art vision-language-action (VLA) models through this lens exposes severe dimension-specific failures invisible to conventional metrics. Through targeted causal intervention, we identify the visual encoder's ability to preserve local spatial structure as a key bottleneck for fine-grained precision: improving it directly unlocks previously inaccessible manipulation capabilities without modifying downstream policies. MetaFine further supports hybrid real-sim validation, using limited paired real-world rollouts to calibrate scalable simulation-based estimates for more stable physical benchmarking. By shifting evaluation from ranking to diagnosis, MetaFine turns benchmarking into an actionable compass for repairing the layered capacities underlying genuine physical dexterity. The MetaFine framework, benchmarks, and supporting resources will be publicly released at our project page: https://metafine.github.io/.

精细操作评估框架具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。