arXiv:2601.18723cs.RO2026-01被引 4

用细粒度评估取代成功率,看清机器人操作的优劣差异

Eval-Actions: Fine-Grained Execution Quality Evaluation for Robotic Manipulation

  • 设计专家评分+动作排序+思维链三重标注体系,精准捕捉执行质量
  • 在1.3万条真实机器人数据上实现90%以上成功识别率,相关性达0.84
  • 适合想提升机器人策略可解释性与诊断能力的研究者使用

尽管视觉-动作(VA)和视觉-语言-动作(VLA)策略推动了机器人操作的发展,其评估仍依赖二元成功/失败率,难以揭示完成相同任务时执行过程的细微差异。本文提出Eval-Actions,一种用于学习型操作策略的细粒度执行质量评估方法及真实机器人基准。该基准包含超过13,000条遥操作与策略生成的实机数据,覆盖150多个任务,约52小时的RGB-D视频、机器人状态轨迹、任务描述及成功/失败标签。其密集标注子集提供专家评分(EG)、基于排序的标签(RG)与思维链式注释(CoT)。我们进一步构建AutoEval,一个从视觉序列与紧凑运动摘要中预测质量评分、任务结果与诊断解释的多模态评估器。在标注测试集上,AutoEval-S在EG与RG下的斯皮尔曼等级相关系数(SRCC)分别为0.81和0.84,成功检测准确率达90.6%与91.0%;AutoEval-P在CoT下达到0.70的SRCC。对专家一致性、物理指标基线、模态消融、结构化泛化与离线策略排序的分析表明,Eval-Actions提供了标准化且可解释的诊断信号,补充传统成功率评估。

原文摘要 · Abstract (English)

Although Vision--Action (VA) and Vision--Language--Action (VLA) policies have advanced robotic manipulation, their evaluation remains dominated by binary success rates, which obscure process-level differences among executions that complete the same task. We introduce Eval-Actions, a diagnostic evaluation methodology and real-robot benchmark for fine-grained execution-quality assessment of learned manipulation policies. Eval-Actions combines criteria-based Expert Grading (EG), Rank-Guided (RG) labels that align measurable motion indicators with expert rankings, and Chain-of-Thought-style (CoT) annotations that explain observable quality differences. The benchmark contains 13K+ teleoperated and policy-generated real-robot episodes covering 150+ tasks and approximately 52 hours of recordings with RGB-D videos, robot-state trajectories, task descriptions, and success/failure labels. Its densely annotated subset provides EG/RG/CoT supervision for training and evaluation. We further provide AutoEval, a reference multimodal evaluator that predicts quality scores, task outcomes, and diagnostic explanations from RGB temporal evidence and compact kinematic summaries. On the annotated Eval-Actions test split, AutoEval-S achieves Spearman rank correlations (SRCCs) of 0.81 and 0.84 under EG and RG, with success detection accuracies of 90.6% and 91.0%; AutoEval-P reaches 0.70 SRCC under CoT. Analyses of expert consistency, physical-metric baselines, modality ablations, structured generalization, and offline policy ranking show that Eval-Actions provides standardized, interpretable diagnostic signals complementary to success-rate evaluation.

机器人操作评估方法细粒度分析真实机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。