提出诊断性微基准,量化视觉智能体的纠错能力与瓶颈。
Evaluating Self-Correcting Vision Agents Through Quantitative and Qualitative Metrics
- 设计微基准,分离任务成功率与纠错成功率。
- 纠错效果在三次尝试后饱和,初始能力不决定修复能力。
- 发现语义漂移是主要失败原因,占比约28%。
多模态基础模型使视觉语言代理(VLAs)能将复杂视觉任务分解为可执行的工具化计划。尽管已有基准开始评估迭代自我修正,但其定量极限和主导推理瓶颈仍不清晰。本文提出诊断性微基准,分析显示任务成功率(TSR = 62%)与纠错成功率(CSR = 25 至 33%)可分离,且初始能力无法预测修复能力。我们明确量化了纠错的边际收益递减现象,其在三次重试后趋于饱和。失败分类表明,约28%的失败源于语义漂移(Semantic Drift),即上下文状态丢失。通过识别这一推理瓶颈,本基准建立了可复现的框架,推动具备状态感知、可信的多模态智能体发展。
原文摘要 · Abstract (English)
Recent progress in multimodal foundation models has enabled Vision-Language Agents (VLAs) to decompose complex visual tasks into executable tool-based plans. While recent benchmarks have begun to evaluate iterative self-correction, its quantitative limits and dominant reasoning bottlenecks remain poorly characterized. This work introduces a Diagnostic Micro-Benchmark. Our analysis decouples Task Success Rate (TSR = 62 percent) from Correction Success Rate (CSR = 25 to 33 percent), revealing that initial competence does not predict repair ability. We explicitly quantify the diminishing returns of correction, which saturates after three retries. Our Failure Taxonomy reveals a frequent factor is Semantic Drift (about 28 percent of failures), a loss of contextual state. By isolating this reasoning bottleneck, this benchmark defines a reproducible framework toward stateful, trustworthy multimodal agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。