评测多模态模型用视觉精细操作工具的能力,发现现有模型普遍表现不佳。
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

- 通过逐步作画还原参考图像,测试模型根据视觉线索精准控制工具
- 当前模型重建相似度仅0.40-0.54,轨迹常提前饱和或后期退化
- 适合关注具身智能、视觉-动作对齐与多模态决策的研究者
评估正从静态问答转向模型通过外部工具行动的代理场景。我们识别出一个关键但未被充分探索的能力——精细视觉工具使用:模型基于视觉证据推断工具参数,且参数直接决定最终结果。现有基准涵盖网页导航、GUI操作和软件工程,但极少关注视觉证据与执行精度之间的耦合。我们提出EASEL,一个评估受控型精细视觉工具使用的基准,以参考引导的视觉重建为主要任务:代理逐步作画以匹配参考图像。EASEL还包含区域标注、手写和路径规划等语义任务。我们进一步提供EASEL-Data,一个包含44万样本的两阶段课程数据集用于轨迹监督,并构建EASEL-9B以研究该数据对能力的影响。对25个模型的评估显示,当前多模态代理在EASEL上系统性表现差:重建相似度瓶颈在0.40–0.54之间,轨迹诊断揭示严重闭环不稳定性——模型通常早期饱和或峰值后退化。语义任务显示在精确标注和路径规划上存在明显能力边界。EASEL-9B在EASEL-Data上训练,相比基础模型相对提升6.3%,在所有评估模型中排名第三。
原文摘要 · Abstract (English)
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual evidence and execution precision. We propose EASEL, a benchmark evaluating a controlled instance of dexterous visual tool use that adopts reference-guided visual reconstruction as its primary proxy task: the agent incrementally paints a canvas to match a reference image. EASEL additionally includes semantic tasks spanning region annotation, handwriting, and path planning. We further provide EASEL-Data, a 440k-sample two-stage curriculum dataset for trajectory supervision, and EASEL-9B to investigate its effect on this capability. Evaluation of 25 models reveals that current multimodal agents systematically struggle on EASEL. Reconstruction similarity bottlenecks at low levels (0.40-0.54), while trajectory diagnostics expose severe closed-loop instability - models typically saturate early or degrade post-peak. Semantic tasks reveal sharp capability boundaries in precision annotation and path planning. EASEL-9B, trained on EASEL-Data, surpasses the base model by a relative 6.3%, ranking third among all evaluated models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。